The Epistemic Operating Model
A framework for trustworthy organizational reasoning across fragmented systems
- Author
- IntraQ
- Published
- Reading time
- 38 minutes

Summary
Organizations spent decades digitizing their work, and the result is a landscape of accurate systems of record that cannot, between them, answer the questions that matter most. Frontier language models make cross-system reasoning practical for the first time, but they introduce probabilistic reasoning into environments where truth, uncertainty, judgment, and authority cannot be allowed to collapse into one another. We argue that the missing layer in enterprise AI is not access, retrieval, context, or tooling, but governed understanding: a discipline about what an organization is justified in believing, what it can currently know, and what it may act on. This paper proposes the Epistemic Operating Model as a framework for handling that tension. The model separates reasoning freedom from claim authority: the reasoning model is free to investigate and interpret, and rules it does not control decide what may be stated as fact, what must be marked as inferred or unknown, and what may be acted on. The paper sets out the problem, the tension at its centre, the distinctions an intelligence system must preserve, and the principles we have arrived at by building and operating such a system in human resources and workforce compliance, one of the most demanding domains available. It proposes an operating model rather than documenting a product, and it describes principles rather than mechanisms.
The problem digitization left behind
The enterprise software era succeeded at its stated goal. Payroll, human resources, benefits, time and attendance, document management, and a hundred narrower functions now run on systems that keep accurate records and enforce their own rules, and a well-run organization can tell you, from the right system, exactly who was paid what, which policy version was published on which date, and which employees acknowledged which handbook. What digitization did not produce is organizational understanding. The questions a leader actually asks rarely live inside one system. Whether a particular state law applies to the organization depends on where people work, which is a payroll fact, on how many of them there are, which is a headcount fact, and on what the law requires at that threshold, which is not an organizational fact at all. Whether the organization is prepared for an audit depends on documents, on evidence of practice, on the population those documents are supposed to cover, and on whether the records describing that population are complete. Each system answers its own question reliably and none of them answers the composite question, because the composite question was never anyone's schema.
For most of the software era this gap was closed by people. An experienced HR leader carries a working model of which system is authoritative for what, which records are stale, which exceptions exist and why, and which questions cannot be answered without a phone call. That model is expensive to build and lost when the person leaves, and the integration and reporting projects that narrowed the gap did so only for the questions that could be anticipated. The questions that matter operationally are often the ones nobody modelled, and they arrive in plain language from someone who needs an answer before a meeting.
Why frontier models change the question
Large language models alter this picture in two directions at once. In one direction they remove the bottleneck. A capable model can read across several systems, notice that a headcount figure and a work-location field together determine whether a threshold has been crossed, and compose an answer that no single system could have produced, for questions that were never modelled, in the vocabulary of the person asking, at a cost that makes routine use plausible. In the other direction they introduce a problem that did not exist while the bottleneck was human. Models are effective reasoners precisely because they infer, generalize, and synthesize beyond what the explicit records say, and that is also why their output cannot be treated as a record. A model that has read a policy document can describe what the policy requires, and it can also describe, with equal fluency, what a similar policy usually requires, what this policy probably intends, and what an employee in this situation would typically be entitled to. In conversation the difference between those statements is often harmless. In an organization, where an answer about eligibility becomes a payroll decision and an answer about applicability becomes a compliance posture, the difference is the whole matter.
The research on model reliability in high-stakes domains is consistent on this point. A Stanford study of leading AI legal research tools found that even retrieval-grounded systems produced incorrect or misgrounded answers a substantial fraction of the time, and that a citation to a real source is not the same as support for the proposition it is cited for (Magesh et al., 2024). Earlier work by the same group measured hallucination rates between 58 and 88 percent for general-purpose models on verifiable questions about federal court cases, and cautioned against rapid deployment of such models in legal work (Dahl et al., 2024). A theoretical account from Kalai and colleagues argues that the procedures models are trained and evaluated under reward guessing over acknowledging uncertainty, so that a model which declines to answer is penalized relative to one that guesses plausibly (Kalai et al., 2025). None of this means models should be excluded from organizational reasoning. It means the architecture around them has to decide, deliberately and outside the model, what may become organizational truth. The model needs freedom to reason, because constrained reasoning is what produced the shallow copilots that users abandon, and the organization needs discipline around what becomes true, because an organization that cannot distinguish its records from its assistant's inferences has lost something it cannot easily recover. The two needs pull in opposite directions only if they are handled by the same mechanism. Our thesis is that they should not be.
The Epistemic Operating Model
An organizational intelligence system becomes useful only when it can investigate across organizational information, reason beyond simple retrieval, and produce useful judgment while preserving explicit boundaries around what is known, what is inferred, what remains uncertain, what is recommended, and what is authorized. Investigation is more than lookup, useful judgment acknowledges that an organization asks for recommendations and not only facts, and the boundaries are what make the rest safe to rely on. We call the operating model that preserves those boundaries the Epistemic Operating Model. The word is chosen for precision rather than weight. Epistemic concerns knowledge: what is known, how it came to be known, what justifies a conclusion, and where the limits of knowing lie. Those are the questions this paper turns out to be about, whether the subject is evidence and fact, fact and interpretation, hypothesis and recommendation, a missing record and a negative one, or a conclusion that was true in March and is being asked about in September. We call it an operating model rather than an architecture because it describes how a system and the people who rely on it conduct themselves, and more than one implementation can honour it.
The operating model rests on one governing doctrine, which we state plainly because everything else follows from it. Code governs truth. The model governs inquiry. Rules written and reviewed by people decide what counts as an organizational fact, which sources are authoritative for which propositions, who may see what, whether a population is complete, whether evidence is current, and whether a statement the model wishes to make is actually supported by the evidence it cites. The model decides what to ask, where to look, which hypothesis to pursue, what the evidence means, and what the organization should consider doing about it. The model is free in the second domain and has no vote in the first.
The consequence of this separation is that every statement the system produces has a type and a status that were assigned outside the model. A fact is a statement established against a governed source under rules the reasoning model does not control. An interpretation is a reading of facts that a person can evaluate. A hypothesis is a possibility stated in order to be tested. A recommendation is judgment resting on facts and readings. An unknown is a statement that something could not be established, together with the governed reason it could not. These types determine what may rest on what, what may be shown as settled, and what must be qualified: a recommendation may rest on facts and readings, a fact may never rest on a recommendation, and a fact that fails its rules is withheld rather than softened. The doctrine also has a second reading, which is that the system should constrain what the model may claim and should not, without necessity, constrain what the model may think about. Early systems built on this doctrine, ours included, over-correct on exactly this point, and we return to that failure below.
A conceptual model
What each layer may produce
Organizational systems
records, each authoritative for its own question
records, not statementsGoverned evidence
what the records establish, with provenance and completeness
KnownIntelligence
investigation across sources; readings and hypotheses
InferredUnknownJudgment
what the organization should consider doing, resting on the above
RecommendedHuman authority
what the organization decides to do
Authorized
Distinctions that decide outcomes
Conversational AI can collapse a great many distinctions without visible cost, because the reader supplies the context and bears the consequences. Organizational intelligence cannot, because the reader is often acting on the answer, and we treat each collapse as a specific defect rather than a general concern about accuracy.
Records, retrieval, and sufficiency
Data is not truth. A record is a claim made by a system at a time, under that system's rules, about some part of the world. Two systems can hold accurate records that disagree, because they observed different things or the same thing at different times, and the disagreement is itself a fact about the organization that should be preserved rather than resolved by whichever system was read last.
Retrieval is not understanding, and relevance is not sufficiency. Finding the records that bear on a question is the beginning of an investigation. A retrieval system can return the three most relevant documents and leave the question unanswerable, because the record that mattered was not indexed or the relevant population is only partly observable. A governed system distinguishes evidence that is present from evidence that is absent, truncated by a limit, or withheld from the current user by authorization, and it reasons differently over each, which is the practical meaning of the knowledge-base literature's point that a system should know where it is incomplete before it reasons over what it has (Razniewski et al., 2024). Missing evidence, in turn, is not negative evidence. The absence of a policy document from the document store is not the absence of a policy, and an employee whose work location is not recorded is not an employee with no work location. Under an open-world reading, "not stated" is distinct from "false", and the correct response to absence is a typed unknown with a governed reason.
The compliance chain
Workforce compliance offers a clean example of a chain of distinctions that must each be preserved. A document is not a policy: a file titled "Leave Policy" may be a draft, a superseded version, a template never adopted, or the current policy, and the store rarely knows which. A policy is not a requirement: the organization's policy may exceed, satisfy, or fall short of what a jurisdiction requires, and the text alone does not say which. A requirement is not applicability: a statute may impose obligations only above a headcount threshold, only in certain states, or only for certain categories of worker, so that whether it applies is a fact composed from several systems. Applicability is not compliance: an obligation that applies may or may not be satisfied, and the evidence of satisfaction lives in acknowledgments, postings, payroll configurations, and practice. A system that lets any link collapse into its neighbour will be wrong in a way that is hard to detect, because each step looks reasonable, and treating a document's existence as evidence of compliance is the most common collapse and the most consequential, since it produces a comfortable answer to the question a leader most wants comforted about.
Change, correlation, and cause
A change is not a cause. Two series that move together in the same quarter have not explained each other, and a turnover increase that coincides with a leadership change has not been caused by it, however natural the narrative. The epidemiological tradition set the standard: Hill's nine viewpoints, temporality among them, are ways to study an association "before we cry causation", and he held that none of them could settle the question on its own (Hill, 1965). A temporal sequence is at most a precondition. A governed system may state that a metric changed, may state that an event occurred in the same period, and may state a causal reading as a hypothesis to be tested. It may not state the cause as a fact, and its rules should refuse causal language in a factual claim as a named defect.
The epistemic ladder
The remaining distinctions form a ladder that every consequential answer climbs. A fact is not an interpretation: that a population's work state is unknown for a quarter of its members is a fact, and that state-level compliance modelling is therefore unreliable is a reading of it, which a person may accept, reject, or qualify. An interpretation is not a hypothesis, since a reading describes what the evidence means while a hypothesis proposes something that further evidence could confirm or refute. A hypothesis is not a recommendation, and a recommendation is not an authorization: the system may recommend that a policy be drafted, and the organization decides whether it will be. Argyris's ladder of inference, developed to describe how people in organizations move from observable data through inferred meanings to beliefs and action, describes a similar ascent and the same danger of acting on rungs whose lower steps were never checked (Argyris, 1982; Argyris, 1990). Two temporal distinctions complete the set and receive their own treatment below: a previous conclusion is not current truth, and a human correction is not automatic truth.
Judgment without counterfeit facts
The most instructive lesson from building a governed system was not about hallucination. It was about usefulness, and it arrived through a production failure we can describe in abstracted form. A system that had correctly established several organizational facts, with sources and verification, was asked which of the resulting issues to address first. No record in the organization contained the answer, because prioritization is not a fact about the world; it is a judgment about what matters given the facts. The system, built to ship only what it could verify, reported that it could not establish a conclusion from verified evidence. Every fact it had gathered was on the page, and the answer to the question was not.
Both obvious corrections are wrong. An architecture that requires the recommendation itself to be an established fact cannot answer any question that requires judgment, and most of the questions a leader asks require judgment. An architecture that lets the model present its judgment as fact restores usefulness at the cost of the distinction that made the system trustworthy, and it does so invisibly, since a confident recommendation reads exactly like a finding.
The resolution is epistemic separation applied to judgment specifically. A recommendation is a first-class kind of statement with its own rules. It must rest on something: on facts the system established, on readings of those facts, on unknowns the system has named. Its status is determined by what it rests on: a recommendation resting only on verified facts is grounded, while one resting on a reading that could not be independently checked, or on something that could not be established, is conditional and is presented as conditional. Its numbers are held to the same standard as a fact's numbers, so that a recommendation cannot smuggle in a figure its basis does not carry. And it is presented as what it is. The answer to "which of these should be fixed first?" reads as a recommendation, with the evidence it rests on stated beside it, so that a reader can see whether they are looking at the organization's records or at the system's judgment about them.
Judgment can then be wrong without any fact being wrong, and when a fact beneath a recommendation changes, the recommendation's status changes with it without anyone having to notice. The system can be as opinionated as a good colleague while remaining as auditable as a ledger, because the two qualities are enforced by different mechanisms.
Unknown and knowability
Every organizational intelligence system will spend much of its time not knowing things, and the difference between a trustworthy system and an untrustworthy one lies largely in how it handles that state. We treat unknown as a first-class organizational state rather than a failure to answer, and we treat it as typed. There is a difference between "we do not know", "the source that would tell us is unavailable", "the evidence is incomplete", "the sources disagree", "the population this would be measured over is only partly observable", "this cannot be established from the records as they stand", and "this was known once and may now be stale". Each has a different governed reason, points to a different remedy, and supports a different kind of further reasoning.
An unknown with a governed reason is evidence: it is a fact about the organization's own observability, and a system that has established that a question cannot currently be answered has established something a leader needs to know. It is also the ground on which honest judgment stands, because a recommendation that rests partly on an unknown is conditional in a way that can be stated.
Knowability is the more demanding idea. An intelligence architecture should reason not only about organizational state but about what can currently be known about that state. The record of an investigation, meaning which sources were read, which were absent, which were truncated by limits, and which the current person was not authorized to see, is itself a first-class result. Two people asking the same question with different authorizations should receive different answers, and each should be told the shape of what they could not see without being told its contents.
Knowability is also the first thing to change. Non-monotonic reasoning teaches that learning something new retracts the statement that it was not known, which means a system's unknowns are the part of its understanding most likely to be stale at any moment. A payroll system that was disconnected yesterday and connected today has changed what the organization can know without changing any fact about the organization, and a governed system that keeps track of its own unknowns can tell a leader which things that could not previously be established are now answerable, without manufacturing a single finding.
Organizational memory is not chat memory
The current generation of AI products has made memory a feature, and most of what is meant by the word is chat memory: recalling what a user said, preferred, or asked for in earlier sessions. Remembering that someone asked about sick leave last week is different from remembering what was established about sick leave, when it was established, which evidence supported it, which reading was made of that evidence, what remained unresolved, and whether the underlying records have changed since. The first is a convenience; the second is the beginning of institutional understanding, and it has properties the first lacks.
The most important is temporal and epistemic boundary. A conclusion reached in March was reached against March's records, by a person with March's authorization, under March's policy versions. When the same question is asked in September, the conclusion is evidence of what was concluded, with a date, and it is not a fact about September. A system that remembers conclusions without their basis and their time will confidently repeat what used to be true, and the failure is worse than having no memory, because it arrives with the authority of prior work. The truth-maintenance literature established decades ago that a system should record the reasons for its beliefs and not merely the beliefs, so that a belief can be explained and retracted when its reasons change (Doyle, 1979), and the W3C provenance model defines derivation and invalidation as separate relations, so that an entity which has expired remains in the record beside what it was derived from rather than being simply removed (W3C, 2013). We would add that a conclusion must keep its clock.
The second property is that organizational memory must be projected through authorization: what one person established with broad access must not become available to another with narrow access merely because both are in the same conversation or the same organization. Memory that leaks through role boundaries is a security defect dressed as a feature, and the research on memory injection in agent systems shows how readily an ungoverned memory becomes an attack surface (Dong et al., 2025; Chen et al., 2024).
The third property, and the heart of the matter, is that conclusions must be reconsiderable rather than merely remembered. A remembered conclusion is a string. A reconsiderable conclusion is a structure: the statement, its type, the evidence it rested on, a way of returning to the sources that evidence came from, and a way of telling whether those sources still say what they said. With that structure a system can notice that a conclusion's evidence has changed, mark the conclusion as superseded rather than repeat it, re-read only the sources behind the premise a person is questioning, and preserve or revise the conclusion for stated reasons. Organizational memory becomes a map of what was established and how, rather than an archive of what was said.
Reasoning continuity
Reconsiderable conclusions make possible a capability we call reasoning continuity, which we distinguish from conversational continuity. Conversational continuity answers "what were we talking about?" Reasoning continuity answers "what did we establish, what supported it, which part of it is this question about, and what needs to be checked again before it can be relied on?"
Consider a system that has investigated an organization and recommended addressing a particular obligation first, and a person who then asks whether the system is sure, or what in the data behind that recommendation should make them doubt it. A system without reasoning continuity treats this as a new question about the organization and investigates from the beginning, rebuilding its picture from scratch and arriving, if it arrives at all, at the same place it started, at several times the cost. We know this failure well because we produced it. A system with reasoning continuity treats the question as being about its own earlier work. It identifies the recommendation being challenged, the facts and readings that recommendation rested on, and the sources those facts came from. It decides, as a matter of judgment, which premise the question is really testing. It re-reads only the sources behind that premise, compares what they say now with what they said then, and preserves, qualifies, or revises the recommendation with an explanation of why.
The discipline is that continuity may not become trust. The earlier conclusion is a map, not an authority. Rules the model does not control decide whether the earlier work belongs to this organization and may be seen by this person, whether its evidence is current or must be re-read, which of its statements were facts and which were readings, and whether anything may be reused as established. An earlier fact whose source now says something different is superseded, and nothing may rest on it. An earlier reading remains a reading and cannot be promoted to fact by the passage of a turn. When the organization has changed since the earlier answer, an earlier fact is possibly stale and must be re-read before a new fact may rest on it. The model decides what is relevant; code decides what is still true.
Reasoning continuity matters for reasons that reinforce one another. Latency and cost fall, because a follow-up that re-reads one source does not pay for an investigation that re-reads a dozen, and the research on long contexts is clear that a model's use of what is in its window degrades as the window grows (Anthropic, 2025a; Chroma, 2025; Liu et al., 2023). Trust rises, because a system that can say which premise it re-checked and what it found behaves like an expert who remembers their own reasoning rather than a search engine that forgets each query. Understanding accumulates, because every follow-up that confirms, supersedes, or qualifies an earlier statement leaves a more precise record than the original investigation did.
Proactive intelligence and the latent question
The strongest organizational intelligence does not always begin with an answer. Sometimes it begins with a question the organization did not know to ask. Two authoritative systems disagree about an employee's status. A conclusion reached in the spring may no longer hold because the organization has changed since. A connected system stopped synchronizing weeks ago, and every conclusion that relied on it has quietly become less reliable. A metric moved during the same period as an operational change.
Each of these is worth a person's attention, and none is a finding. A latent question is not a conclusion. An intelligence system can surface something worth investigating without manufacturing an answer, and the discipline lies in keeping the two apart. The last example is the clearest test: a metric that changed during the same period as an event is a coincidence in time until something more is established, and a system that presents it as a cause has crossed from noticing into fabrication. A governed system presents it as a question, names the two observations it rests on, and lets the investigation that follows establish whatever can be established.
The site reliability literature offers the operational discipline. Alerting should catch symptoms rather than causes, except for very definite and imminent ones, and every page should be actionable, require intelligence, and concern a novel problem (Beyer et al., 2016). An alerting strategy is judged on precision, recall, detection time, and reset time together, and alerts that keep firing after a problem is resolved lead to confusion or to issues being ignored (Beyer et al., 2018). Base rates sharpen the point, since even an accurate detector applied to a rare condition produces mostly false alarms, which is why clinical decision-support alerts are overridden at the high rates the medical informatics literature has documented for two decades (van der Sijs et al., 2006), and why "question" is the only honest framing for what an organizational detector finds. Proactive intelligence, properly governed, is therefore a modest and demanding thing. It surfaces latent questions from governed state, ranks them by how much they matter, carries the basis for each so that a person can see why it was raised, and never lets one become a finding by any path other than an investigation. It is at its most valuable when it points at the organization's own observability: the source that stopped reporting, the population that cannot be seen, the conclusion whose ground has shifted.
General intelligence and organizational intelligence together
A governed system that only conducts investigations will be consulted rarely, because most of what people do all day does not require an investigation. The research on how people use AI assistants is clear that daily habit forms at the trivial end of the range. Anthropic's first Economic Index report found a slight lean toward augmentation, with 57 percent of observed tasks augmented rather than automated, augmentation meaning that people iterate with the model, learn from it, or validate their own work with it (Anthropic, 2025b), and the NBER study of ChatGPT usage found that "asking" interactions, meaning requests for information and advice, are both the largest category and the most highly rated by users (Chatterji et al., 2025). The thirty-second task and the three-day problem are not competing priorities; the first is how the second earns its place.
The thirty-second task in HR is usually a matter of general professional knowledge. How should a manager structure a performance conversation? What should a job description contain? These questions do not depend on the organization's records, and a system that responds by assembling an organizational picture and launching an investigation is both slow and slightly absurd. The answer should come from general professional knowledge, in one round, labelled as such, offering the organizational follow-up rather than manufacturing one. The next question often does depend on the records: what happened with this employee previously? That is governed organizational context, and the answer must come from authoritative sources, under the current person's authorization, with the completeness of those sources stated. The question after that may need both: given that history, how should the manager prepare for tomorrow's conversation? Here general knowledge about how such conversations go must work together with organizational facts about this person, and the architecture's task is to let them work together without letting one become the other.
General model knowledge must not silently become organizational truth. A statement drawn from general knowledge is labelled as general, and it may not carry an organization-specific value or status, nor a figure that no governed source supplies. In our experience the most dangerous default in enterprise AI products is a silent fallback from grounded retrieval to general knowledge without saying that the mode has changed, and a product that blends a knowledge base, a general model, and customer data without saying which answered leaves the reader unable to tell which is which. The legal research finding that a retrieval-grounded tool can cite a real source for an answer that is nonetheless wrong, so that the citation serves only to mislead, is the same failure seen from the other side (Magesh et al., 2024). A governed system keeps the boundary visible in the answer itself, so that "from your records", "the system's judgment", "general professional guidance", and "not established" read differently and mean different things.
Human control and legitimate authority
Human control is usually presented as a safety wrapper, a confirmation placed between an autonomous system and an action to catch mistakes. That framing undersells the organizational reason for it and produces designs that treat approval as friction to be minimized. Organizations have legitimate authority structures. Who may publish a policy, adopt a document, approve a termination, change a classification, or commit the organization to a position is a matter of governance, delegation, and accountability, decided before any AI system arrived. A system that investigates, explains, drafts, and recommends is doing valuable work, and none of it confers authority to publish, adopt, discipline, terminate, approve, or change policy.
This is why recommendation is not authorization, and why the distinction is structural rather than procedural. A recommendation is a statement of judgment, resting on evidence, offered to the people who hold authority. An action is a change in the organization's state, taken by someone who holds the authority to take it, against a target the organization can identify unambiguously, under conditions that hold at the moment of action rather than at the moment of recommendation, and with its effect confirmed rather than assumed. Operational platforms have converged on a similar shape from the other direction: Palantir Foundry's action types, for example, define declared, typed edits to objects, property values, and links, and gate submission with criteria whose conditions each carry a failure message explaining why the user is blocked (Palantir, n.d.). The system's role is to make the people who hold authority better informed, not to hold authority itself, and the architecture should make it impossible for a recommendation, however confident, to become an action by any route that bypasses the organization's own structure.
Compounding organizational understanding
The hardest question any enterprise AI system must answer is what it would possess, after years of use, that a competitor connecting an equally capable model to the same systems tomorrow would not. If the answer is "the model" or "the integrations", there is no durable position, because both are available to the newcomer on the same terms.
Our answer is that the durable asset is governed understanding accumulated through use. It consists of organizational history: dated conclusions, each with the evidence it rested on and the reading that was made of it. It consists of evidence relationships, of unresolved questions and the reasons they could not be resolved, of prior investigations that can be continued rather than repeated, of conflicts between sources preserved rather than resolved, of decisions taken and declined, of corrections and disputes with their classifications, and, over time, of the organization's own vocabulary and the institutional expertise people externalize through their interactions with the system. None of this can be obtained from a connector. A connector delivers the current state of a system of record; systems of record overwrite, and few can answer what the organization knew on the date it made a decision. A newcomer with the same model and the same connectors starts with zero conclusions, zero preserved conflicts, zero recorded decisions, and zero corrections. The asset compounds because its rate of accumulation is bounded by use and by human review, which is exactly why it cannot be bought.
Compounding understanding is only an asset if it does not compound error, and the same distinctions that govern a single answer govern the archive. A correction is not truth: it is a recorded dispute with a classification and a reviewer. A previous conclusion is not current: it carries a date and a basis, and it is superseded when its basis changes. A historical decision does not prove the decision was right: it proves that it was made, by whom, on what evidence, and a system that treats past decisions as precedents to follow rather than records to consult has confused organizational memory with organizational authority. Provenance, time, and type are what let accumulated understanding create value without accumulating error, and the test of a memory architecture is whether its oldest entries can still be read for what they are.
Institutional expertise
People hold knowledge that systems do not. They know why an exception exists, why a process works differently in practice than in the document, what a term means locally, which source is unreliable for which purpose, and why an apparently reasonable conclusion is wrong in this organization. An intelligence system that cannot learn from those people will plateau at whatever its records support, and one that learns from them naively will be poisoned by the first confident mistake.
The naive design is to remember what the person said, and the literature on organizational learning loops shows both why that is tempting and why it fails. Meta's account of an internal system that learns from expert corrections is instructive for its discipline: corrections are diagnosed by attribution into distinct classes before anything changes, every change lands as a human-reviewed proposal after adversarial review and regression testing, and the regression suite grows with each fix (Meta Engineering, 2026). The contrast is with memory systems in which the model itself decides, for each newly extracted fact, whether to add a memory, update one, delete one, or do nothing, deleting memories that new information contradicts, a design adequate for user preferences and unacceptable for organizational facts (Chhikara et al., 2025). Community fact-checking learned a parallel lesson, surfacing a note only when raters who usually disagree agree about it, because a single confident voice is not evidence of truth (Wojcik et al., 2022).
The correct framing of the problem is therefore not "remember what the human said" but "understand what kind of correction occurred and what, if anything, should change". A correction may indicate an actual error in a source, in which case the system's fact was faithfully wrong. It may indicate incomplete evidence, a reasoning failure, an entity-resolution error in which two records were wrongly joined or separated, outdated information, a difference in terminology, an organizational position that no record captures, or a genuine ambiguity. Each class calls for a different response, and only some call for anything to become true. A correction that is a position is a legitimate and valuable thing to record, but it is a position, attributed and reviewed, and it sits beside the general rule rather than overriding it.
We regard this as a frontier rather than a solved problem. What a governed system can do today is capture a correction as a first-class dispute, attributed to a person and a role, classified against a closed vocabulary, carried into subsequent reasoning as contested rather than as truth, and reviewed by a person before it changes anything. What remains to be built is the machinery by which reviewed corrections become durable, scoped, dated, expiring organizational positions that a system can rely on with the discipline it applies to records, together with the replay and regression infrastructure that Meta's account rightly treats as non-negotiable. The principle guiding the work is already clear: memory is context, and enforcement is code. A correction may change what is displayed and what is asked next; it may not change what the rules will pass.
Dependency through usefulness
The word dependency is used in two senses that must not be confused. One is manipulative: a product made hard to leave through artificial switching costs, variable rewards, and engagement mechanics. The economics of switching costs distinguishes learning and transaction costs, which reflect real costs of switching, from artificial or contractual costs that arise entirely at a firm's discretion and carry no such cost (Klemperer, 1987). We are not describing the manipulative kind, and we would not build it.
The other is dependency through usefulness, which generative AI has already demonstrated at the level of the individual. People who can write emails still ask an assistant to help; people who understand spreadsheets still ask for the formula; people who can summarize a document still have it summarized. The NBER study argues that the economic value of a general assistant comes primarily from decision support, with people using it as an advisor or research assistant rather than only as a technology that performs tasks directly (Chatterji et al., 2025). Organizational intelligence may produce the same shift for operational work. The desired instinct is not that a person cannot answer the question without the system; it is that asking is the easiest way to get a governed answer, labelled, sourced, and continuous with what was established before. The thirty-second question builds the habit and populates the record, the three-day problem is where the record pays for itself, and dependency of the legitimate kind persists only as long as the system remains the easiest place to think about and do a category of work.
One intelligence, many systems
There is a recurring temptation to recreate organizational fragmentation at the layer that was supposed to overcome it, by deploying a swarm of specialized agents, one for payroll, one for policy, one for compliance, one for people, that hand each other messages. We make no absolute claim against multi-agent architectures; there are workloads, particularly parallel ones over many independent items, where fan-out is the right design, and Anthropic's account of its own multi-agent research system shows the pattern working (Anthropic, 2025c). The research is more cautious than the enthusiasm, however. Anthropic's guidance on building agents recommends the simplest design that works and adding orchestration only when measurement shows the need (Anthropic, 2024).
The organizational reason is more fundamental than cost. Dependent reasoning across system boundaries needs a coherent reasoning context. Whether a state law applies is a question that crosses payroll, headcount, and the requirements catalog in a single inference; the fact from one system is a premise in a chain whose conclusion lives in another, and the provenance of that conclusion is the whole chain. Split the chain across agents that each see one system, and either the provenance fractures at every hand-off or an orchestrator must reconstruct it, at which point the orchestrator is the intelligence and the agents are connectors. The systems underneath can and should remain specialized; that is what systems of record are for. The intelligence above them should reason across the organization as one context, with one doctrine about truth, one vocabulary of statement types, one memory, and one set of authorization boundaries. We call this, conceptually, one intelligence, and we regard it as a consequence of the doctrine rather than a preference about topology.
HR as the proving ground
We built our first system on this operating model for human resources and workforce compliance, and the choice was fortunate rather than incidental. HR is unusually demanding along every axis that matters to the architecture. Its systems are fragmented by design, its facts change constantly and matter at specific times, and jurisdiction cuts through everything, with obligations differing by state, by city, by headcount, and by classification. Evidence is incomplete in ways the records do not announce, records from different sources conflict about the same person, and decisions are consequential for individuals, the organization, and regulators. Human judgment is irreducible, since no record contains the answer to how a difficult conversation should go, and authorization boundaries are sensitive, because the same system holds compensation, health, and disciplinary information. The work is high-frequency and trivial for most of the day and low-frequency and grave when it is not, so the same system must serve both.
A domain with all of those properties punishes every collapsed distinction described in this paper, quickly and visibly, which is what makes it a proving ground, and an architecture that preserves the distinctions under those conditions has learned something general, because the properties are not unique to HR. Operations faces fragmented systems, temporal change, incomplete evidence, and consequential decisions in its own vocabulary of assets, vendors, incidents, and obligations, and finance faces them in the vocabulary of ledgers, controls, close processes, and filings. Neither is the subject of a product announcement here; we are describing an architectural implication. The underlying problem is organizational rather than uniquely HR-specific, and a reasoning layer that answers it well in one domain is built from parts that do not know which domain they are in.
What the Epistemic Operating Model is not
The Epistemic Operating Model is adjacent to several technologies, each of which solves an important part of the problem and none of which solves the whole of it. Enterprise search and retrieval-augmented generation solve access and relevance: they find the records that bear on a question and put them in front of a model, but they do not decide what is authoritative for what, whether the evidence is sufficient, whether a population is complete, or whether the model's synthesis is supported by what it cites. A chatbot over company documents inherits the same limits and adds the collapse of document into policy. A tool-calling protocol such as MCP solves connectivity and is agnostic by design about truth. A data warehouse solves integration and historical query for the questions modelled in advance, and carries no epistemic type. A knowledge graph gives structure to entities and relations, but a graph populated by a model is only as trustworthy as the model's extractions, and a graph does not by itself distinguish a fact from an inference or record why an edge was retracted. Deterministic expert systems held the line on truth and could not reason beyond their rules; autonomous agents optimize for task completion on the basis of whatever they believed; traditional business intelligence answers the questions its dashboards were built to answer; and generic copilots are used episodically because they lack context, forget, and give general answers to specific questions.
A system built on the operating model uses retrieval, tools, integration, structure, and models, and adds the layer that decides what the organization is justified in believing as a result: typed statements with provenance and status, verification the reasoning model does not perform on itself, typed unknowns with governed reasons, memory that carries time and basis, judgment separated from fact, and action separated from recommendation. That layer is the missing one, and it is not produced by improving any of the others.
Architectural principles
We can now state the principles this paper has argued for. They are offered as a synthesis rather than a checklist, and each follows from the doctrine that opened the paper.
-
Code governs truth; the model governs inquiry. Rules written and reviewed by people decide what counts as a fact, which sources are authoritative, who may see what, and whether a statement is supported. The model decides what to ask, where to look, and what the evidence means.
-
Reasoning freedom and claim authority are separate mechanisms. The system constrains what the model may claim and does not, without necessity, constrain what it may think about. Every statement carries a type and a status assigned outside the model, and the types determine what may rest on what.
-
Unknown is a first-class organizational state, and knowability is part of the state. An unknown carries the governed reason it could not be established. The record of what was read, absent, truncated, or unauthorized is itself a result, and universal claims are bounded by the coverage of the population they are made over.
-
Judgment rests on evidence without becoming evidence. A recommendation is grounded when it rests on verified facts and conditional otherwise, and it is presented as judgment. Judgment can be wrong without any fact being wrong, and it changes status when its basis changes.
-
Organizational memory preserves time, provenance, and authorization. A conclusion keeps its date, its basis, and the role that reached it. What was concluded is evidence of what was concluded, never a fact about the present.
-
Previous reasoning is reconsiderable, not merely remembered. A follow-up about earlier work continues from its basis, re-reads selectively, and preserves or revises for stated reasons. Changed evidence supersedes what rested on it; nothing earlier is trusted by the passage of time.
-
Latent questions precede unsupported conclusions. The system surfaces what is worth investigating, with its basis, and manufactures no findings. Symptoms are noticed; causes are established or stated as hypotheses.
-
General knowledge and organizational evidence work together without merging. General knowledge is labelled and never carries an organizational fact; organizational evidence is sourced and never inferred. The boundary is visible in the answer.
-
Human correction improves understanding without silently rewriting truth. A correction is a recorded, attributed, classified dispute, reviewed by a person. Memory is context; enforcement is code.
-
Consequential action requires legitimate organizational authority. A recommendation is not an authorization. Actions are typed transitions on identified targets, taken by people who hold the authority, confirmed afterwards.
-
The intelligence layer becomes simpler to use as the architecture beneath it becomes more sophisticated. The person asks a question and receives an answer with a stance. Everything that makes the answer trustworthy happens beneath the surface, and none of it is the person's burden to understand.
Conclusion
The enterprise does not lack data, and it does not lack access to models that can reason over data. What it lacks is a layer that can say, with discipline, what the organization is justified in believing, what it can currently know, and what it may act on, while remaining useful for the small questions that fill a working day and the large ones that define a quarter. We have described the operating model for that layer, which we call the Epistemic Operating Model, in terms of the problems it must solve and the distinctions it must keep, rather than the mechanisms by which it keeps them. The doctrine at its centre is simple to state and demanding to honour: code governs truth, and the model governs inquiry.
The next generation of enterprise software will not be defined only by what systems can record or by what models can generate. It will be defined by what organizations can reliably understand. The Epistemic Operating Model is our attempt to define how AI can participate in that understanding without erasing the boundaries between evidence, inference, uncertainty, judgment, and authority.
References
- Anthropic (2024). Building Effective Agents. Anthropic Engineering, December 2024.
- Anthropic (2025a). Effective Context Engineering for AI Agents. Anthropic Engineering, September 2025.
- Anthropic (2025b). The Anthropic Economic Index. Anthropic, February 2025.
- Anthropic (2025c). How We Built Our Multi-Agent Research System. Anthropic Engineering, June 2025.
- Argyris, C. (1982). Reasoning, Learning, and Action: Individual and Organizational. Jossey-Bass.
- Argyris, C. (1990). Overcoming Organizational Defenses: Facilitating Organizational Learning. Allyn and Bacon.
- Beyer, B., Jones, C., Petoff, J., and Murphy, N. R., eds. (2016). Site Reliability Engineering: How Google Runs Production Systems, chapter 6, "Monitoring Distributed Systems" (written by Rob Ewaschuk). O'Reilly Media.
- Beyer, B., Murphy, N. R., Rensin, D. K., Kawahara, K., and Thorne, S., eds. (2018). The Site Reliability Workbook, chapter 5, "Alerting on SLOs". O'Reilly Media.
- Chatterji, A., Cunningham, T., Deming, D. J., Hitzig, Z., Ong, C., Shan, C. Y., and Wadman, K. (2025). How People Use ChatGPT. NBER Working Paper 34255, September 2025.
- Chen, Z., Xiang, Z., Xiao, C., Song, D., and Li, B. (2024). AgentPoison: Red-teaming LLM Agents via Poisoning Memory or Knowledge Bases. Advances in Neural Information Processing Systems 37 (NeurIPS 2024).
- Chhikara, P., Khant, D., Aryan, S., Singh, T., and Yadav, D. (2025). Mem0: Building Production-Ready AI Agents with Scalable Long-Term Memory. arXiv:2504.19413.
- Chroma (2025). Context Rot: How Increasing Input Tokens Impacts LLM Performance. Hong, K., Troynikov, A., and Huber, J. Chroma Research, July 2025.
- Dahl, M., Magesh, V., Suzgun, M., and Ho, D. E. (2024). Large Legal Fictions: Profiling Legal Hallucinations in Large Language Models. Journal of Legal Analysis, 16(1), 64–93.
- Dong, S., Xu, S., He, P., Li, Y., Tang, J., Liu, T., Liu, H., and Xiang, Z. (2025). A Practical Memory Injection Attack against LLM Agents. arXiv:2503.03704, version 1, March 2025.
- Doyle, J. (1979). A Truth Maintenance System. Artificial Intelligence, 12(3), 231–272.
- Hill, A. B. (1965). The Environment and Disease: Association or Causation? Proceedings of the Royal Society of Medicine, 58(5), 295–300.
- Kalai, A. T., Nachum, O., Vempala, S. S., and Zhang, E. (2025). Why Language Models Hallucinate. arXiv:2509.04664.
- Klemperer, P. (1987). Markets with Consumer Switching Costs. The Quarterly Journal of Economics, 102(2), 375–394.
- Liu, N. F., Lin, K., Hewitt, J., Paranjape, A., Bevilacqua, M., Petroni, F., and Liang, P. (2023). Lost in the Middle: How Language Models Use Long Contexts. arXiv:2307.03172; Transactions of the Association for Computational Linguistics.
- Magesh, V., Surani, F., Dahl, M., Suzgun, M., Manning, C. D., and Ho, D. E. (2024). Hallucination-Free? Assessing the Reliability of Leading AI Legal Research Tools. arXiv:2405.20362; published in the Journal of Empirical Legal Studies, 22(2), 2025.
- Meta Engineering (2026). An Organizational Second Brain: Building an AI That Learns From Experts. Engineering at Meta, September 2026.
- Palantir (n.d.). Action Types: Overview. Palantir Foundry Documentation. Accessed September 2026.
- Razniewski, S., Arnaout, H., Ghosh, S., and Suchanek, F. (2024). Completeness, Recall, and Negation in Open-World Knowledge Bases: A Survey. ACM Computing Surveys, 56(6), Article 150.
- van der Sijs, H., Aarts, J., Vulto, A., and Berg, M. (2006). Overriding of Drug Safety Alerts in Computerized Physician Order Entry. Journal of the American Medical Informatics Association, 13(2), 138–147.
- W3C (2013). PROV-DM: The PROV Data Model. Moreau, L., and Missier, P., eds. W3C Recommendation, April 2013.
- Wojcik, S., Hilgard, S., Judd, N., Mocanu, D., Ragain, S., Hunzaker, M. B. F., Coleman, K., and Baxter, J. (2022). Birdwatch: Crowd Wisdom and Bridging Algorithms can Inform Understanding and Reduce the Spread of Misinformation. arXiv:2210.15723.