← All writing

Essay

Truth and Inference

How philosophies of truth help us judge what to expect from LLM inference and test its outcomes in practice.

The most dangerous prompt is often the one that already sounds solved. Consider a company that imports customer spreadsheets. One request says:

Our customers keep choosing the wrong columns. Design an AI mapper that fixes their mistakes automatically and reduces support costs.

Another says:

Customers submit files with unfamiliar headers, optional identifiers, mixed date formats, and fields whose meanings vary by account. We have the target schema, representative files, validation errors, and examples of manual corrections. Propose candidate mappings, expose ambiguous interpretations, and identify the evidence needed before any mapping is applied automatically.

The first request is clean and may be right, but it smuggles in three untested claims: customers are making mistakes, automatic mapping is the right intervention, and fewer support tickets would prove success. A model can polish those claims without discovering whether they describe the problem. The less elegant second request separates observation from interpretation, names its evidence, allows several causes, and gives the answer somewhere to fail.

Core question

How do we recognize meaningful input and valuable output when AI can make a weak premise and a strong answer sound equally plausible?

Philosophies of truth give us questions to ask of an answer. LLM inference produces predictions shaped by context; information theory helps explain how that context changes what the model is likely to produce. The practical bridge is to make the outcome we expect explicit, then test the result: does it fit its assumptions, match what we observe, and hold up when someone uses it?

Meaningful input exposes evidence, assumptions, constraints, omissions, and uncertainty. Valuable output survives checks against the intended outcome. Understanding and Bottlenecks put judgment in bounded teams. This essay connects philosophies of truth to the everyday work of those teams, before The Knowledge Factory turns that practice into a repeatable system.

01

The Input Is Part of the Argument

We often evaluate AI after it answers, checking sources, logic, code, and conclusions. That starts too late. A prompt argues by naming the problem, selecting observations, excluding explanations, and implying what to optimize. Wrong choices make a faithful response more convincingly wrong.

“Customers keep choosing the wrong columns” sounds observational but may be interpretive: perhaps the interface hides definitions, teams use the same header differently, the schema lacks a needed distinction, or uncertain mappings are applied too confidently. A model cannot recover possibilities the task treats as settled; it reasons from the world the prompt supplies.

Good input enables discrimination: what supports or contradicts the diagnosis, which terms are shared, where uncertainty remains, and who resolves it? Ask which output claims follow from premises, depend on outside facts, can be tested before action, or emerge in use. A request is meaningful when others can inspect its frame. Neither prompt nor response may grade itself.

Prompt quality

Detail can still be nonsense. Improve the epistemics: connect the request to evidence, expose assumptions to challenge, and preserve unresolved questions for the right person or test.

02

A Good Answer Can Be Wrong in Three Ways

Three overlapping truth practices catch different failures: coherence, correspondence, and consequence.

Coherence asks: Does it fit? A signup_date → createdAt mapping may fit the schema, types, and other mappings. That supports internal consistency, not that the source means the date the application expects.

The philosophical theory is more demanding and contested: philosophers dispute both coherence and its relevant set of beliefs or propositions. Incompatible claims can fit incomplete accounts, so consistency alone is too weak. Here, coherence tests rather than exhausts truth.

Correspondence asks: Does it match? Documentation, files, customer explanations, and application behavior may reveal that signup_date means an export date, account opening, or administrator entry. The name is evidence, not terrain.

Consequence asks: What happens when someone relies on it? A mapping may fit the schema and sample yet discard timezones, merge customer concepts, or cost more to clean up than manual entry. Consequences do not make temporary or commercial success true or good; they reveal what abstractions do to people and systems over time.

Three tests

Coherence · Correspondence · Consequence

Formal truth

Coherence — Does it fit?

validity relative to definitions, axioms, and inference rules

Language favors — explicit premises, symbolic relationships, proof obligations

Feedback — counterexamples and proof assistants reject invalid derivations

Empirical truth

Correspondence — Does it match?

agreement with an observable state of affairs—the events, objects, properties, or relations the claim describes

Language favors — measurement, method, uncertainty, replication, counterevidence

Feedback — failed predictions and unreplicated results erode the claim

Operational truth

Consequence — Does it work?

reliable consequences under stated conditions—the procedure repeatedly produces its intended result within defined tolerances

Language favors — procedures, preconditions, failure modes, tolerances, observed outcomes

Feedback — systems that crash, stall, or cost too much are corrected or retired

Internal consistency, external observation, and reliance answer different questions. Passing one test does not guarantee passing the next.

The order varies: observation can expose contradiction, a bad consequence can reveal a missing observation, and formal inconsistency can stop action early. Success under one test must not stand in for all three.

Other practices matter: acquaintance tests fidelity to experience, sincerity aligns expression with belief, and trustworthiness asks whether reliance is earned. Vision and Values develops them for values, testimony, and relationships; here the three technical and product tests recur.

Two theological parallels help situate these non-propositional practices without centering them. Confucian chéng joins freedom from deceit with inward-outward integrity; biblical Hebrew ʾemet can mean truth, faithfulness, firmness, and reliability. Neither tradition is equivalent to this framework.

Practice changes language: definition, evidence, and correction refine an idea; continuity, clarification, and specification discipline it into a term of art—a stable label for needed distinctions.

AI Factory · 02 · Refinement + discipline

Refinement and discipline establish a term of art

Scroll the path →

Refinement and discipline establish a term of artRefinement defines and corrects a concrete idea; discipline sustains continuity, clarification, and specification until it becomes a term of art with a stable domain boundary.Disciplinecontinuity · clarificationspecificationRefinementdefinition · evidence · correctionIdeashared meaningTerm of artconsistent within a domain
Refinement supplies definition, evidence, and correction. Discipline supplies continuity, clarification, and specification. Together they give an idea a stable expression as a term of art.

A term of art is a compact address into a history of distinctions, not a guarantee. What happens when that address enters a predictive model?

03

The Right Term Changes the Prediction

J. R. Firth's “You shall know a word by the company it keeps” captured Zellig Harris's related distributional idea: use patterns reveal linguistic structure. Word2Vec represented similarly used words with related vectors. Modern models are neither giant Word2Vec tables nor complete distributional theories of meaning, but this lineage explains how terms direct prediction without containing answers.

The semantic-composition developer tool lets you combine up to four terms and inspect their three-dimensional projection and vector arithmetic. It illustrates distributional geometry, not measured semantic identities.

AI Factory · 02b · Embedding space / grounded inference

A grounded term can guide useful expansion

Scroll the path →

A grounded term can guide useful expansionA short label used as a grounded term of art guides inference toward expanded text that fits the intended result.01 / WITH UNDERSTANDINGShort labelshared term of artEmbedding spaceUnderstandingTerm of artshared domain distinctionsInferenceguided by meaningUseful expansionFits the intended resultmore text · useful tokens
A term of art packs shared domain distinctions into a short label. When those distinctions are learned or supplied in context, they can guide model representations and inference toward useful expanded output. This is a conceptual path, not a measured gain: embeddings represent input; the model generates text, and value must be checked against the intended result.

A tokenizer assigns text tokens identifiers; an embedding table maps them to learned vectors. Position, attention, and feed-forward layers produce contextual states, then an output projection and softmax produce next-token probabilities.

token IDs → input embeddings → contextual hidden states → output logits → next-token probabilities

Context changes the state: stable participates in different relationships in stable counting sort and stable employment. A domain term selects among patterns associated during training or supplied in the prompt; it does not store them verbatim.

Information theory describes that selection. Surprisal rises as an observed outcome's assigned probability falls. Entropy, expected surprisal across a distribution, is higher when uncertainty is spread broadly than when concentrated on a few outcomes. Conditional prediction asks how context changes that distribution.

Shannon developed entropy for communication theory. In 1951, next-character guesses in unfamiliar English estimated linguistic redundancy. Neural language models do not directly implement that experiment but share its problem: assign probabilities to what follows. During training, cross-entropy penalizes too little probability on observed continuations, teaching conditional regularities rather than fixed phrase retrieval.

Compare these requests:

Map the date column.

Map activation_date to a nullable calendar-date field. Reject timestamps and locale-ambiguous values; expose rows that require account-specific review.

The second request's type, rejected cases, and review boundary narrow competent continuations without making the model generally smarter; this is not a measured probability change. More distinctions aid prediction and evaluation, yet the result can be precisely wrong. If activation_date records an export job, a type-safe explanation still maps the wrong concept. Conditional probability shows how an answer follows from context, not whether context matches the world.

Developer note — entropy, surprisal, and cross-entropy

For an outcome x with probability p(x), surprisal is commonly written I(x) = -log₂ p(x). Shannon entropy is expected surprisal: H(P) = -Σ p(x) log₂ p(x). Cross-entropy scores a predicted distribution against observations from another distribution. In next-token training, the loss decreases as the model assigns more probability to the observed token.

These equations describe distributions. They do not measure truth, meaning, or business value.

Prediction explains a term's efficiency; practice determines whether it was applied well.

04

Domains Teach Language What to Reject

Technical language carries constraint history: terms name needed distinctions, syntax records relationships, examples establish ordinary cases, failures draw boundaries, and tools, institutions, and consequences reinforce surviving uses.

A term of art therefore points beyond its glossary entry into a bounded context of assumptions, operations, and known failures. The same process preserves stale habits, blind spots, and fashionable mistakes: language remembers repetition, not only truth.

Label → operationalize → formalize → compute

A label names a distinction; practice connects it to action and consequence; formalization makes its relationships composable; computation derives, executes, and checks some of them. Each step sharpens the next but must return to observation.

Working hypothesis

From labeling the world to computation

Read downward. Correspondence grounds labels; consequence operationalizes and compresses those that work; coherence formalizes their relations; computation applies the resulting structure.

  1. Labeling

    Correspondence · Does it match?

    Labels are tested against the world

    Observation, measurement, and counterexamples correct names that fail to track events, objects, properties, or relations.

    observe + name + correct
  2. Operationalization

    Consequence · Does it work?

    Useful distinctions become efficient terms

    Repeated practice compresses successful inputs, operations, boundaries, and failure modes into terms of art that guide action.

    execute + select + compress
  3. Formalization

    Coherence · Does it fit?

    Explicit rules make labels compositional

    Mathematics, logic, type systems, and programming languages specify how symbols may combine and what follows.

    abstract + relate + derive

Computation

Resulting capability — not a fourth theory of truth

Formal relations become machine-operable

Machines can generate synthetic cases, type-check programs, prove derivations, simulate models, and execute tests.

generate + derive + execute

Regrounding closes the loop. A proof establishes derivability from stated premises; a program may compile and pass its tests. Neither result alone shows that the premises model the world, the synthetic data are representative, or the outcome is worth pursuing.

Hypothesis — correspondence grounds labels; consequence selects and compresses those that work; coherence formalizes their relations. Computation uses that structure, but its results must be regrounded.

Domains expose different errors: mathematics uses definitions, proof obligations, and counterexamples; empirical inquiry uses measurement, replication, and failed prediction; software uses specifications, types, tests, runtime behavior, and operational signals; product work adds customer evidence, adoption, correction effort, cost, and human effects. No master score subsumes them.

Code makes some constraints executable. Hash-based sorting sounds specific but names no universal algorithm: what is hashed, which key ranges matter, how collisions behave, whether order is stable, what complexity is acceptable, and what invalid input does all remain open.

Once those assumptions are explicit, the work can face several independent checks:

  • A type checker rejects relationships that violate declared types.
  • Unit examples establish named behavior.
  • Property tests challenge broader invariants such as sorted order and preservation of elements.
  • Benchmarks compare latency or memory under a stated workload.
  • Runtime observation catches conditions the test environment omitted.
  • A product owner may still reject the implementation because it solves the wrong problem.
Corpuscontains the patterns accumulates what practitioners wrote under real conditions
Syntaxrejects invalid token sequences before anything else runs
Typesrejects invalid relationships between values and operations
Testsrejects specified behavioral failures
Runtimerejects crashes, latency, and resource misuse
Usersrejects behavior that fails in the world — physical, economic, human
The code constraint stack — every layer filters invalid expressions, which is why the language that survives is unusually pattern-dense.

Some bounded coding tasks approach evaluative closure: the repository contains enough evidence, constraints, feedback, and delegated authority to judge improvement.

inspect → change → compile → test → benchmark → compare → revise

Strategy rarely closes so neatly: evidence may be tacit, private, contested, or emerging; effects arrive late; environments react; and people dispute “better.” Tests encode acceptance criteria, not whether a goal deserves pursuit.

Vibes to a Typographic Specification

For the THOM logo, broad prompts produced plausible but uncontrollable marks. Domain language—optical profiles, stroke hierarchy, glyph silhouette, spacing, construction-line density—specified Bézier curves, stroke thickness, and the header/footer's compact M. You can probably tell it was vibe designed, but observable constraints improved it. This proves no general theory; it offers a task-level diagnostic:

  • Is the vocabulary stable inside a bounded context, or are key terms overloaded?
  • Do examples include counterexamples and boundary cases, or only clean successes?
  • Can an external check reject plausible output, or does evaluation depend on how it sounds?
  • Does the task preserve uncertainty and disagreement, or force everything into one confident answer?

A model can be sharp where feedback is dense and glib where it is thin. Valid TypeScript under tests says little about an ambiguous customer field.

05

Paid Work Begins After Generation

Return to the import-mapping product. This is a composite example, not a report of measured revenue, accuracy, or customer savings; it exposes the chain a real product must close.

The system needs a target schema, representative files, field definitions, aliases, accepted mappings, corrections, boundary cases, the customer's aim, and a decision boundary separating suggestions, automatic application, and mandatory review.

The model proposes candidates and their interpretations: company might mean an account, employer, legal entity, or free text; active might mean current state, historical observation, or billing flag. Hiding alternatives improves readability while reducing trust.

Mechanical checks narrow the space: types reject impossible values, parsers catch invalid formats, referential constraints expose missing identifiers, and required fields, enumerations, and uniqueness rules reject other candidates. They establish schema coherence and anticipate some consequences, but cannot settle a customer's ambiguous source meaning.

A domain owner then contributes evidence the system lacks: field use, reversible loss, affected records, and when uncertainty must stop the process. This is decision-making, not ceremonial approval after the model decides.

Only use reveals value. Observe accepted imports, corrections, reversals, support work, and customer time saved; count generation, review, integration, monitoring, failure, and support costs. More output is not more value. A technically valid mapping that creates expensive cleanup is a bad product result.

StageWhat must remain visibleWhat can reject the work
Supplied contextDefinitions, representative data, aims, and known ambiguityMissing or contradictory evidence
Candidate mappingAssumptions and alternative interpretationsSchema, parser, and referential checks
Domain reviewWorkflow meaning, reversibility, and affected recordsAn authorized owner
Customer resultCorrections, reversals, support work, and time savedUnacceptable consequence or total cost
Retained learningRevised definitions, boundary cases, and stopping rulesNew evidence from the next use

Commercial value appears after candidates survive evaluation, produce an acceptable result, and justify total delivery cost. This describes a mechanism, not measured ROI. Meaningful input exposes enough of the problem to distinguish candidates; valuable output carries assumptions and uncertainty into a process that can check them. Failure should add corrected definitions, boundary examples, or stopping rules to maintained context. Durable value appears when evaluation changes what the organization knows, not merely when generation produces another candidate.

06

Make Judgment Local and Inspectable

One skilled reviewer can check a handful of decisions, but accelerated generation turns that expert into a queue. Scale requires no lower standard; it requires a usable local one: shared definitions, evidence requirements, acceptance criteria, tools exposing contradictions, failed tests, uncertainty, and downstream effects, plus decision rights for local conclusions, cross-team commitments, and escalation.

A compact claim record can keep the reasoning visible:

FieldQuestion
ClaimWhat are we asking others to rely on?
SourceWhich observation, document, test, or person supports it?
AssumptionsWhat must be true for the claim to hold?
ScopeWhere does it apply, and where does it stop?
UncertaintyWhich interpretations or outcomes remain unresolved?
DisconfirmationWhat evidence would weaken or overturn it?
AuthorityWho may act, approve, stop, or escalate?
ConsequenceWhat happened after someone relied on it?

The table does not create knowledge. It attaches a conclusion to the conditions under which it earned reliance, letting another team challenge its premise without reconstructing the conversation and preserving disagreement long enough to become useful.

Preserve counterexamples and failed consequences, not only approvals. Recording success while erasing correction teaches people and models the wrong lesson. Local results that contradict a shared assumption must change the context; otherwise alignment is merely a coherent map nobody may compare with terrain.

The practice

Inspect the input → generate candidates → discriminate among claims → act within the evidence → observe consequences → revise the shared context.

AI can surface premises, retrieve sources, generate alternatives, execute tests, compare outcomes, and record change. It should not collapse those steps into one fluent answer and call the work complete. Meaningful input makes prediction's grounds inspectable; valuable output survives coherence, correspondence, and consequence against an explicit aim.

A model is fluent where language has learned to carry the constraints. Our work is to follow that fluency back through consequence to correspondence: does it fit, does it work, and does it match the world we mean to change?

The next article, The Knowledge Factory, builds a repeatable system around these teams and their standards of judgment: how does intent become a verified outcome?

07

Sources