[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"sanity-dAaBBNIyXv6Yhuq_jCGgGAKjH4c4syA_c1D-ksFovGo":3,"sanity-wE9gZprPI1yUpfIGus9Q7SQDjldY7AkXc-efsUjAdqU":1041},{"data":4,"sourceMap":-1},{"latestPodcast":5,"latestReleases":14,"post":39,"recent":1016},[6],{"_id":7,"publishedAt":8,"slug":9,"sponsored":12,"title":13},"8e6c3a8a-3d27-44c8-be90-0b74d1090a4a","2026-10-06T17:00:00.000Z",{"_type":10,"current":11},"slug","tales-from-the-2026-developer-survey-results",null,"Tales from the 2026 Developer Survey results",[15,21,27,33],{"_id":16,"publishedAt":17,"slug":18,"title":20},"7a2d88e5-ee53-4f7c-a46e-47bed1cebadf","2026-09-30T16:00:00.000Z",{"_type":10,"current":19},"anyone-can-start-building-verified-knowledge-with-stack-internal","Anyone can start building verified knowledge with Stack Internal",{"_id":22,"publishedAt":23,"slug":24,"title":26},"12c6a8a7-135f-401f-ac1a-1c26c33e69c0","2026-09-03T16:00:00.000Z",{"_type":10,"current":25},"security-control-and-accessibility-si-2026-6","Elevating security, control, and accessibility: Stack Internal 2026.6",{"_id":28,"publishedAt":29,"slug":30,"title":32},"adcf1bca-3295-4ac5-9b3c-23f337974190","2026-07-30T15:10:00.000Z",{"_type":10,"current":31},"introducing-stack-internal-new-platform-experience","Your trusted knowledge layer: Introducing Stack Internal's new platform experience",{"_id":34,"publishedAt":35,"slug":36,"title":38},"eb5b66eb-9410-4329-83bb-22bbff39402a","2026-04-28T13:00:00.000Z",{"_type":10,"current":37},"turn-scattered-knowledge-into-trusted-intelligence","Turning scattered knowledge into trusted intelligence: Stack Internal 2026.3",{"_createdAt":40,"_id":41,"_rev":42,"_system":43,"_type":46,"_updatedAt":47,"author":48,"body":59,"comments":926,"dateUrl":927,"image":928,"product":12,"publishedAt":934,"seo":935,"slug":938,"sponsored":12,"tags":940,"title":1015,"visible":926},"2026-10-07T20:14:22Z","39f2fc9f-742a-4273-a4d1-abdecd14fcf8","46s78gX2DRxVswzX0l1bjy",{"base":44},{"id":41,"rev":45},"9fZfA2bz90M2vN1VVZmTvQ","blogPost","2026-10-07T20:38:29Z",[49],{"_createdAt":50,"_id":51,"_rev":52,"_type":53,"_updatedAt":54,"employee":55,"name":56,"slug":57},"2026-10-07T20:10:11Z","c7434ebf-211a-49d5-a564-22a91d63ed1e","9fZfA2bz90M2vN1VVZP5Ov","blogAuthor","2026-10-07T20:12:01Z","none","Varun Jindal",{"_type":10,"current":58},"varun-jindal",[60,76,84,109,118,150,155,167,170,178,200,228,240,256,264,288,291,294,302,310,318,330,333,353,356,368,372,388,391,407,415,439,455,458,461,469,472,488,528,536,552,555,607,615,631,635,651,659,682,746,749,765,768,784,792,812,824,836,848,867,887,895,910,918],{"_key":61,"_type":62,"children":63,"markDefs":74,"style":75},"f71e6b664971","block",[64,69],{"_key":65,"_type":66,"marks":67,"text":68},"e2e628f53b0b","span",[],"The trick to putting LLM agents in high-stakes systems isn’t a smarter model — it’s containing the model to one node so the rest of the system is ordinary, testable code. Here are the structural moves, with the contracts ",{"_key":70,"_type":66,"marks":71,"text":73},"3c8adc404e1e",[72],"em","and types to implement them.",[],"normal",{"_key":77,"_type":62,"children":78,"markDefs":83,"style":75},"7392fd9e6d11",[79],{"_key":80,"_type":66,"marks":81,"text":82},"685f197747cd",[],"Demos love autonomous agents that loop, call tools, and “figure it out.” Production hates them. The moment an agent’s behavior depends on which path the model wandered down today, you can’t test it, can’t audit it, and can’t let it touch anything that matters. In a regulated or high-consequence system — money movement, healthcare, infrastructure — “it usually works” is a non-starter.",[],{"_key":85,"_type":62,"children":86,"markDefs":108,"style":75},"472271c0070e",[87,91,96,100,104],{"_key":88,"_type":66,"marks":89,"text":90},"7f02a79a5d26",[],"This is Level 1 of a six-level maturity model for running LLM systems in production: ",{"_key":92,"_type":66,"marks":93,"text":95},"b821d00d8711",[94],"strong","the determinism layer",{"_key":97,"_type":66,"marks":98,"text":99},"a80b911f0a5f",[],". Before you can do evals, confidence routing, or anything else higher up the stack, you need the part underneath to hold still. The whole game at this level is one idea — ",{"_key":101,"_type":66,"marks":102,"text":103},"3182d8942fcd",[94],"contain the model to a single node so everything around it is ordinary, testable code",{"_key":105,"_type":66,"marks":106,"text":107},"48d974d6e802",[]," — expressed through a handful of design moves. You keep the model’s intelligence and remove almost all of the unpredictability. This post gives the contracts, not just the concepts.",[],{"_key":110,"_type":62,"children":111,"markDefs":116,"style":117},"1a857388bd11",[112],{"_key":113,"_type":66,"marks":114,"text":115},"89c8776e440d",[],"Agents propose; they don’t act",[],"h3",{"_key":119,"_type":62,"children":120,"markDefs":149,"style":75},"38eebe8fafd6",[121,125,129,133,137,141,145],{"_key":122,"_type":66,"marks":123,"text":124},"2b334608ed94",[],"The most important rule: ",{"_key":126,"_type":66,"marks":127,"text":128},"21c8a862fde1",[94],"an agent is a pure function",{"_key":130,"_type":66,"marks":131,"text":132},"e04067c7f867",[]," from context to a ",{"_key":134,"_type":66,"marks":135,"text":136},"96c40a6e2078",[72],"proposed",{"_key":138,"_type":66,"marks":139,"text":140},"349b69528bc1",[]," decision — pure with respect to business state, and deterministic given its gateway (the one nondeterministic call, the model, is injected so tests can fake it). It has no authority to change the world. A separate, dumb, heavily-tested component — the ",{"_key":142,"_type":66,"marks":143,"text":144},"c22f0965ad12",[94],"substrate",{"_key":146,"_type":66,"marks":147,"text":148},"3d7e195b8fd6",[]," — applies proposals, and only after the required approval.",[],{"_key":151,"_type":152,"code":153,"language":154,"markDefs":12},"69f422a004cf","code","from typing import Protocol, Literal\nfrom dataclasses import dataclass\n\n@dataclass(frozen=True)\nclass Proposal:\n    decision_id: str\n    capability: str\n    action: dict                  # the structured, proposed change — NOT yet applied\n    confidence: float\n    routing: Literal[\"auto\", \"hitl_recommended\", \"hitl_required\", \"reject\"]\n    reasoning: list[str]\n    evidence: list[dict]\n\nclass Agent(Protocol):\n    def propose(self, ctx: \"Context\") -> Proposal: ...     # pure: no side effects on business state\n\nclass Substrate(Protocol):\n    def apply(self, proposal: Proposal, approval: \"Approval\") -> \"Effect\": ...  # the ONLY mutator","python",{"_key":156,"_type":62,"children":157,"markDefs":166,"style":75},"4d4ac66ca665",[158,162],{"_key":159,"_type":66,"marks":160,"text":161},"7d61db390749",[],"Because ",{"_key":163,"_type":66,"marks":164,"text":165},"2d36ced467ba",[152],"propose() is pure, its headline property is one assertion:",[],{"_key":168,"_type":152,"code":169,"language":154,"markDefs":12},"0c836ade0717","def test_propose_is_pure():\n    agent = ClassifyAgent(gateway=FakeGateway(scripted))\n    assert agent.propose(ctx) == agent.propose(ctx)   # same context → same proposal; no world touched",{"_key":171,"_type":62,"children":172,"markDefs":177,"style":75},"e332a66eeb7e",[173],{"_key":174,"_type":66,"marks":175,"text":176},"92fe8edd4279",[],"The boundary buys you three properties:",[],{"_key":179,"_type":62,"children":180,"level":197,"listItem":198,"markDefs":199,"style":75},"0f4895c2e699",[181,185,189,193],{"_key":182,"_type":66,"marks":183,"text":184},"b7031142b0da",[94],"Testable.",{"_key":186,"_type":66,"marks":187,"text":188},"964fc5439b1e",[]," ",{"_key":190,"_type":66,"marks":191,"text":192},"830edc0bdd9d",[152],"propose()",{"_key":194,"_type":66,"marks":195,"text":196},"85b29609e7ad",[]," is a pure function — same context, same proposal. No mocking the world to test the logic.",1,"bullet",[],{"_key":201,"_type":62,"children":202,"level":197,"listItem":198,"markDefs":227,"style":75},"ea463287d6a7",[203,207,211,215,219,223],{"_key":204,"_type":66,"marks":205,"text":206},"7ce3bc26bfa7",[94],"Safe.",{"_key":208,"_type":66,"marks":209,"text":210},"2477f0e024e4",[]," A jailbroken or buggy agent produces a bad ",{"_key":212,"_type":66,"marks":213,"text":214},"c920d68c1d6c",[72],"proposal",{"_key":216,"_type":66,"marks":217,"text":218},"af305ca65899",[],", not a bad ",{"_key":220,"_type":66,"marks":221,"text":222},"ee30cd74fa8f",[72],"action",{"_key":224,"_type":66,"marks":225,"text":226},"009d6f8455f0",[],". The blast radius stops at “a guardrail or a human said no.”",[],{"_key":229,"_type":62,"children":230,"level":197,"listItem":198,"markDefs":239,"style":75},"7250194a3b29",[231,235],{"_key":232,"_type":66,"marks":233,"text":234},"e4d94d1903e2",[94],"Composable.",{"_key":236,"_type":66,"marks":237,"text":238},"c59a89760761",[]," Agents never call other agents. Work flows through the substrate (e.g. a cases table + a scheduler), so there are no hidden chains of side effects to reason about.",[],{"_key":241,"_type":62,"children":242,"markDefs":255,"style":75},"c3a1045aaddd",[243,247,251],{"_key":244,"_type":66,"marks":245,"text":246},"93f91ece1d43",[],"This single constraint turns “an AI did something we can’t explain” into “an AI ",{"_key":248,"_type":66,"marks":249,"text":250},"0bf5cf7c54fe",[72],"suggested",{"_key":252,"_type":66,"marks":253,"text":254},"6d032970ab58",[]," something, and here’s exactly what approved it.”",[],{"_key":257,"_type":62,"children":258,"markDefs":263,"style":117},"4a0f5df1099d",[259],{"_key":260,"_type":66,"marks":261,"text":262},"221ceaf3d867",[],"One fixed graph per capability",[],{"_key":265,"_type":62,"children":266,"markDefs":287,"style":75},"3af3393cf0ae",[267,271,275,279,283],{"_key":268,"_type":66,"marks":269,"text":270},"83132bd2dc39",[],"Free-form ReAct loops are great for exploration and terrible for guarantees. Model each capability as a ",{"_key":272,"_type":66,"marks":273,"text":274},"1bfae7cc3a5a",[94],"fixed sequence of nodes",{"_key":276,"_type":66,"marks":277,"text":278},"4226f7ccdb4f",[]," where the model is used only where judgment is needed and everything else is ordinary code. The arrow sketch below is abridged for readability; the full node list (with ",{"_key":280,"_type":66,"marks":281,"text":282},"1f13d57a323e",[152],"pre_check and ",{"_key":284,"_type":66,"marks":285,"text":286},"a04a30797a16",[],"memory_write`) is in the next section.",[],{"_key":289,"_type":152,"code":290,"language":154,"markDefs":12},"fe1cd2142938","entry → load context → reason (LLM) → output guardrail → verify →\n        judge (sampled) → compose confidence → route → prepare proposal → record → exit",{"_key":292,"_type":152,"code":293,"markDefs":12},"0f2fbcd5b90d","                Free-form ReAct loop     Fixed-graph (this)\n--------------  -----------------------  ---------------------------------\nControl flow    model decides next step  known in advance\nTestable        hard (path varies)       each node in isolation\nLatency \u002F cost  unbounded                bounded, predictable\nAudit           reconstruct from trace   uniform row every time\nRight for       open-ended exploration   decisions that must be guaranteed",{"_key":295,"_type":62,"children":296,"markDefs":301,"style":75},"98bdcf503d9c",[297],{"_key":298,"_type":66,"marks":299,"text":300},"ac5dc44cbe85",[],"You give up some cleverness; you get back the ability to reason about what the system will do. The clever part — judgment on messy inputs — stays exactly where the model is good at it, in one node, surrounded by code you can read.",[],{"_key":303,"_type":62,"children":304,"markDefs":309,"style":117},"df3931929bfa",[305],{"_key":306,"_type":66,"marks":307,"text":308},"3649adb1e787",[],"Anatomy of a request",[],{"_key":311,"_type":62,"children":312,"markDefs":317,"style":75},"f4749d4aaf10",[313],{"_key":314,"_type":66,"marks":315,"text":316},"826c3b5b0c7d",[],"“Make agents deterministic” is hard to act on until you’ve seen the shape of one. So let’s walk a single request through the graph, node by node — the implementation behind the diagram above.",[],{"_key":319,"_type":62,"children":320,"markDefs":329,"style":75},"32cca6e038ca",[321,325],{"_key":322,"_type":66,"marks":323,"text":324},"d6b5c987b446",[94],"The state object.",{"_key":326,"_type":66,"marks":327,"text":328},"1efe92819ec7",[]," Every node reads and writes one typed state value threaded through the graph. Making this explicit is half the battle — it’s what lets you test a node in isolation by constructing a state and asserting on the result.",[],{"_key":331,"_type":152,"code":332,"language":154,"markDefs":12},"171eb0d29ee9","from typing import TypedDict, Literal, Optional\n\nclass GraphState(TypedDict):\n    decision_id: str\n    tenant_id: str\n    identity: dict                      # who\u002Fwhat authority (set at entry)\n    inputs: dict                        # validated request\n    context: dict                       # loaded memory\u002Freference slices\n    model_output: Optional[dict]        # raw structured output from the LLM node\n    guardrail: dict                     # blocks\u002Fredactions applied\n    verification: dict                  # deterministic check results\n    judge: Optional[dict]               # second-opinion result (if sampled)\n    confidence: Optional[float]         # composed score\n    routing: Optional[Literal[\"auto\", \"hitl_recommended\", \"hitl_required\", \"reject\"]]\n    proposal: Optional[dict]            # final shaped proposal\n    ledger_entry_id: Optional[str]",{"_key":334,"_type":62,"children":335,"markDefs":352,"style":75},"8105199ed65f",[336,340,344,348],{"_key":337,"_type":66,"marks":338,"text":339},"13ce822a69d2",[94],"The node contract.",{"_key":341,"_type":66,"marks":342,"text":343},"4af89ee8c033",[]," Every node is the same shape: ",{"_key":345,"_type":66,"marks":346,"text":347},"34de367c0498",[152],"state -> state",{"_key":349,"_type":66,"marks":350,"text":351},"37d06ba27f0b",[],". Deterministic except the one model node. This uniformity is why you can unit-test each node and reason about the whole.",[],{"_key":354,"_type":152,"code":355,"language":154,"markDefs":12},"3618959efa22","from typing import Protocol\n\nclass Node(Protocol):\n    name: str\n    def run(self, state: GraphState, deps: \"Deps\") -> GraphState: ...",{"_key":357,"_type":62,"children":358,"markDefs":367,"style":75},"648b023859d8",[359,363],{"_key":360,"_type":66,"marks":361,"text":362},"08194aa5e140",[94],"The graph.",{"_key":364,"_type":66,"marks":365,"text":366},"e7660462cc9c",[]," Eight nodes are shared across every capability; only a few are capability-specific. New capability = implement ~4 nodes, inherit the other 8.",[],{"_key":369,"_type":152,"code":370,"language":371,"markDefs":12},"664fd6993615","1  entry              mint decision_id, bind tenant + identity              [shared]\n2  pre_check          validate inputs, resolve refs, cheap early-outs       [capability]\n3  context_load       fetch only the context this decision needs            [shared]\n4  llm_decision       the reasoning step — structured in, structured out    [capability]\n5  output_guardrail   PII \u002F policy scrub on the model output                [shared]\n6  verification       deterministic, capability-specific correctness checks [capability]\n7  judge              sampled second-model review (high-stakes)             [shared]\n8  confidence_compose composed score from the signals                       [shared]\n9  routing            auto vs hitl vs reject                                [shared]\n10 prepare_proposal   shape the final proposal                              [capability]\n11 memory_write       write the agent's own audit memory (not business state)[shared]\n12 exit               append the immutable ledger entry                     [shared]","text",{"_key":373,"_type":62,"children":374,"markDefs":387,"style":75},"eabc1ad70652",[375,379,383],{"_key":376,"_type":66,"marks":377,"text":378},"9bea97c90ca2",[],"The two load-bearing nodes are ",{"_key":380,"_type":66,"marks":381,"text":382},"9efbd890d33b",[152],"llm_decision and ",{"_key":384,"_type":66,"marks":385,"text":386},"5f49a8897a81",[],"verification`. The first is the only nondeterministic node — structured input, schema-constrained output, retry-on-mismatch — and everything around it treats its output as untrusted until checked (more on that next section). The second is deterministic, capability-specific code that the whole “contain the model” thesis rests on:",[],{"_key":389,"_type":152,"code":390,"language":154,"markDefs":12},"fdd62f2afa41","def run(self, state, deps):\n    out = state[\"model_output\"]\n    checks = {\n        \"in_enum\":   out[\"decision\"] in ALLOWED_DECISIONS,        # can't return an off-list action\n        \"schema_ok\": matches_schema(out, DECISION_SCHEMA),\n        \"rules_ok\":  deps.rules.check(out, state[\"inputs\"]),      # business invariants\n    }\n    state[\"verification\"] = {\"checks\": checks, \"score\": sum(checks.values()) \u002F len(checks)}\n    return state",{"_key":392,"_type":62,"children":393,"markDefs":406,"style":75},"0493ffb019fb",[394,398,402],{"_key":395,"_type":66,"marks":396,"text":397},"dd4973499831",[],"The point of the fixed shape: you can read the control flow (the path ",{"_key":399,"_type":66,"marks":400,"text":401},"562bd6ecceb1",[72],"is",{"_key":403,"_type":66,"marks":404,"text":405},"db2aa4fd43c0",[]," the graph), the model stays contained to one node, and every decision is uniform — same shape every time means the same audit row every time, and the same place to add a check, a metric, or a guardrail.",[],{"_key":408,"_type":62,"children":409,"markDefs":414,"style":117},"940d188991a0",[410],{"_key":411,"_type":66,"marks":412,"text":413},"fbe30af3ee95",[],"Structured output over free-form",[],{"_key":416,"_type":62,"children":417,"markDefs":438,"style":75},"8421ceb8e59b",[418,422,426,430,434],{"_key":419,"_type":66,"marks":420,"text":421},"a2123d13e0fc",[],"Node 4 deserves its own treatment, because a surprising amount of LLM fragility comes from one choice: letting the model return free text and then parsing it. Prose is ambiguous, the format drifts between calls, and your downstream code is one unexpected phrasing away from breaking. “Sure! It’s probably Approve, though it could be Escalate” — now you own an NLU problem to extract ",{"_key":423,"_type":66,"marks":424,"text":425},"a323e3f7f103",[152],"Approve",{"_key":427,"_type":66,"marks":428,"text":429},"8b329270b661",[],", and tomorrow's rephrase breaks your regex. You've coupled your system to the model's ",{"_key":431,"_type":66,"marks":432,"text":433},"6b99f9c73d42",[72],"prose style",{"_key":435,"_type":66,"marks":436,"text":437},"f2d3833bbe15",[],", the least stable thing about it.",[],{"_key":440,"_type":62,"children":441,"markDefs":454,"style":75},"a08db6ab494e",[442,446,450],{"_key":443,"_type":66,"marks":444,"text":445},"d67eeffb8cef",[],"The fix is boring and powerful: ",{"_key":447,"_type":66,"marks":448,"text":449},"d986fe826ba4",[94],"constrain the model to emit a validated structure",{"_key":451,"_type":66,"marks":452,"text":453},"da998b476cd6",[],", and treat anything else as a failed call to retry. There are three ways to constrain, picked by how hard the guarantee must be:",[],{"_key":456,"_type":152,"code":457,"markDefs":12},"f0d01d2e5835","method                                         how                             guarantee                             use when\n---------------------------------------------  ------------------------------  ------------------------------------  --------------------------------------\nSchema-guided (JSON Schema \u002F response_format)  ask for JSON matching a schema  strong, provider-enforced             most cases\nTool \u002F function call                           model emits a typed tool call   strong; natural for \"do X with args\"  the decision maps to an action\nGrammar-constrained decoding                   constrain tokens to a grammar   hard guarantee (can't emit invalid)   strict\u002Fregulated formats, local models",{"_key":459,"_type":152,"code":460,"language":154,"markDefs":12},"3666d69cddec","# the contract as types — downstream consumes a typed object, never prose\nfrom pydantic import BaseModel\nfrom typing import Literal\n\nclass Decision(BaseModel):\n    decision: Literal[\"approve\", \"escalate\", \"reject\"]   # an enum is itself a guardrail\n    confidence: float\n    reasons: list[str]",{"_key":462,"_type":62,"children":463,"markDefs":468,"style":75},"6a3dd9d05565",[464],{"_key":465,"_type":66,"marks":466,"text":467},"323f4da8c50d",[],"Constrained generation reduces malformed output; it doesn’t eliminate it. So close the loop: validate every response, and on a mismatch retry — feeding the validation error back. Bound the retries and fail closed.",[],{"_key":470,"_type":152,"code":471,"language":154,"markDefs":12},"e7c374a8a908","from pydantic import ValidationError\nclass NonConformingOutput(Exception): ...\n\ndef decide(base_prompt, model, max_retries=2) -> Decision:\n    prompt = base_prompt\n    for _ in range(max_retries + 1):\n        raw = model.generate(prompt, schema=Decision.model_json_schema())\n        try:\n            return Decision.model_validate_json(raw)            # success: typed object\n        except ValidationError as e:\n            # rebuild from base_prompt — don't append onto the already-appended prompt (it compounds)\n            prompt = f\"{base_prompt}\\n\\nYour previous output was invalid: {e}. Return JSON only.\"\n    raise NonConformingOutput()      # fail closed — never hand downstream a guess",{"_key":473,"_type":62,"children":474,"markDefs":487,"style":75},"5f5f28c6eaec",[475,479,483],{"_key":476,"_type":66,"marks":477,"text":478},"5219974e1093",[],"Validation at the boundary means the rest of the system only ever sees well-formed decisions; the messy “did the model behave” question is contained to this one function. You trade a vague NLU problem for a crisp validation problem — no parsing layer, stability across model\u002Fprompt swaps, a built-in guardrail (a fixed enum ",{"_key":480,"_type":66,"marks":481,"text":482},"45a370b4346e",[72],"cannot",{"_key":484,"_type":66,"marks":485,"text":486},"115afad241e4",[]," return something off-list), and a typed contract golden tests can assert against.",[],{"_key":489,"_type":62,"children":490,"markDefs":527,"style":75},"e519e96e1c4e",[491,495,499,503,507,511,515,519,523],{"_key":492,"_type":66,"marks":493,"text":494},"30bd8a0a6ad0",[],"One caveat worth stating loudly: structure constrains ",{"_key":496,"_type":66,"marks":497,"text":498},"1bd37162bb06",[72],"form",{"_key":500,"_type":66,"marks":501,"text":502},"dea7b1b8304e",[],", not ",{"_key":504,"_type":66,"marks":505,"text":506},"0bc1156e4d1f",[72],"correctness",{"_key":508,"_type":66,"marks":509,"text":510},"cde419be704a",[],". A perfectly valid ",{"_key":512,"_type":66,"marks":513,"text":514},"185241cca38e",[152],"{\"decision\":\"approve\",\"confidence\":0.99}",{"_key":516,"_type":66,"marks":517,"text":518},"732a4817fe67",[]," can be completely wrong. Structured output removes the parsing failure mode, not the judgment failure mode — which is why it's the ",{"_key":520,"_type":66,"marks":521,"text":522},"c1e2dfda0b2f",[72],"floor",{"_key":524,"_type":66,"marks":525,"text":526},"704245792579",[]," of a production system, under evals, verification, and confidence, not a substitute for them.",[],{"_key":529,"_type":62,"children":530,"markDefs":535,"style":117},"65c5d339e67a",[531],{"_key":532,"_type":66,"marks":533,"text":534},"dc701f244978",[],"The glue: confidence → routing",[],{"_key":537,"_type":62,"children":538,"markDefs":551,"style":75},"dd7c2c31f915",[539,543,547],{"_key":540,"_type":66,"marks":541,"text":542},"f70726ea3242",[],"Confidence (model signal + verification + sampled judge) is composed at the confidence node; the routing node turns it into one of four paths with a single threshold, and ",{"_key":544,"_type":66,"marks":545,"text":546},"8fdeac2850f8",[152],"prepare_proposal",{"_key":548,"_type":66,"marks":549,"text":550},"cfcdfc65503b",[]," later stamps that result onto the proposal:",[],{"_key":553,"_type":152,"code":554,"language":154,"markDefs":12},"740b102b7b10","def route(confidence: float, verified: bool, T: float = 0.85) -> str:  # T is per-capability config, not a constant\n    if not verified:           return \"reject\"\n    if confidence >= T:        return \"auto\"\n    if confidence >= T - 0.2:  return \"hitl_recommended\"   # close: pre-fill the proposal for a human\n    return \"hitl_required\"                                 # low: a human decides from scratch",{"_key":556,"_type":62,"children":557,"markDefs":606,"style":75},"0ddc956ae3f6",[558,562,566,570,574,578,582,586,590,594,598,602],{"_key":559,"_type":66,"marks":560,"text":561},"407005641a94",[152],"verified",{"_key":563,"_type":66,"marks":564,"text":565},"9d6534a99141",[]," is derived from the verification node (",{"_key":567,"_type":66,"marks":568,"text":569},"2ebc43aad1b1",[152],"verified = all(checks.values())",{"_key":571,"_type":66,"marks":572,"text":573},"fec22596def8",[],"), not a field the agent sets on itself, and the router emits all four ",{"_key":575,"_type":66,"marks":576,"text":577},"780a66347894",[152],"routing",{"_key":579,"_type":66,"marks":580,"text":581},"a20ce45bce06",[]," states. Note the order: routing runs ",{"_key":583,"_type":66,"marks":584,"text":585},"e60debf83038",[72],"before",{"_key":587,"_type":66,"marks":588,"text":589},"e9bdd279fee0",[]," the proposal is shaped (node 9 → node 10), so it takes a plain confidence and a verified flag — not a ",{"_key":591,"_type":66,"marks":592,"text":593},"0d91ccf48168",[152],"Proposal",{"_key":595,"_type":66,"marks":596,"text":597},"87b3de5ffdd8",[],". Start ",{"_key":599,"_type":66,"marks":600,"text":601},"babbaf14e3f0",[152],"T",{"_key":603,"_type":66,"marks":604,"text":605},"2b1f9d4fa486",[]," conservative — everything to a human — and lower it per slice only as data proves it safe. Human attention, the expensive resource, gets spent exactly where the system is unsure.",[],{"_key":608,"_type":62,"children":609,"markDefs":614,"style":117},"da65b707cd9a",[610],{"_key":611,"_type":66,"marks":612,"text":613},"962e8c6525f0",[],"The ledger: an append-only record of record",[],{"_key":616,"_type":62,"children":617,"markDefs":630,"style":75},"45f9bdbaee25",[618,622,626],{"_key":619,"_type":66,"marks":620,"text":621},"a3fec0a84356",[],"Every proposed decision becomes ",{"_key":623,"_type":66,"marks":624,"text":625},"3abdd272e7b0",[94],"one immutable row",{"_key":627,"_type":66,"marks":628,"text":629},"ffc821269b07",[]," — written as the last node of every decision, unconditionally. Not a log line; the canonical record of what happened and why.",[],{"_key":632,"_type":152,"code":633,"language":634,"markDefs":12},"e018cbeed906","CREATE TABLE decision_ledger (\n  decision_id     TEXT PRIMARY KEY,\n  ts              TIMESTAMPTZ NOT NULL,\n  tenant_id       TEXT NOT NULL,\n  capability      TEXT NOT NULL,\n  inputs_hash     TEXT NOT NULL,          -- hash, not the raw sensitive payload\n  model_version   TEXT NOT NULL,\n  prompt_version  TEXT NOT NULL,\n  decision        JSONB NOT NULL,\n  confidence      REAL NOT NULL,\n  routing         TEXT NOT NULL,          -- auto | hitl_* | reject\n  outcome         TEXT,                    -- recorded later as a NEW superseding row, never an in-place UPDATE\n  supersedes      TEXT REFERENCES decision_ledger(decision_id),\n  prev_hash       TEXT,                    -- optional hash-chain for tamper-evidence\n  entry_hash      TEXT\n);\n-- append-only: no UPDATE\u002FDELETE grants; corrections are new rows that set `supersedes`.","sql",{"_key":636,"_type":62,"children":637,"markDefs":650,"style":75},"c59552946939",[638,642,646],{"_key":639,"_type":66,"marks":640,"text":641},"df4dfa618cfb",[],"Hash the inputs, don’t warehouse them — verifiability without the liability. Make it append-only: corrections supersede, never overwrite. If you can ",{"_key":643,"_type":66,"marks":644,"text":645},"70bb25151bc4",[152],"UPDATE",{"_key":647,"_type":66,"marks":648,"text":649},"ccc308610705",[]," the ledger, it’s not an audit trail; revoke the grant. And never skip the write under load — it’s the record of record, not droppable telemetry.",[],{"_key":652,"_type":62,"children":653,"markDefs":658,"style":117},"c51eeb6cda87",[654],{"_key":655,"_type":66,"marks":656,"text":657},"039a47c68bf6",[],"Bounded ReAct: the one place loops belong",[],{"_key":660,"_type":62,"children":661,"markDefs":681,"style":75},"1a89d4e953b6",[662,666,669,673,677],{"_key":663,"_type":66,"marks":664,"text":665},"0c486976948b",[],"If you’ve followed all of the above — fixed graphs, the model contained to one node — you’ll eventually hit a problem that doesn’t fit: something open-ended where the model must look something up, reason about what it found, maybe look up more, then decide. That ",{"_key":667,"_type":66,"marks":668,"text":401},"c704d151dce8",[72],{"_key":670,"_type":66,"marks":671,"text":672},"ee8f904ab80f",[]," what ReAct-style tool loops are for. The mistake isn’t the loop; it’s the ",{"_key":674,"_type":66,"marks":675,"text":676},"5c0991cf9e46",[72],"unbounded",{"_key":678,"_type":66,"marks":679,"text":680},"3870395fd5b6",[]," loop. Allow autonomy as a deliberate, rail-guarded exception — and nowhere else.",[],{"_key":683,"_type":62,"children":684,"markDefs":745,"style":75},"e779c474d038",[685,689,693,697,701,705,709,713,717,721,725,729,733,737,741],{"_key":686,"_type":66,"marks":687,"text":688},"6879f782debe",[],"Four rails keep it controlled: ",{"_key":690,"_type":66,"marks":691,"text":692},"ac2ec5f2594f",[94],"(1) a hard iteration cap",{"_key":694,"_type":66,"marks":695,"text":696},"eb714a0fac17",[]," — ",{"_key":698,"_type":66,"marks":699,"text":700},"e8e43847b96a",[152],"for, never ",{"_key":702,"_type":66,"marks":703,"text":704},"2530684c357b",[],"while(not done)`, so the model doesn’t decide when to stop; ",{"_key":706,"_type":66,"marks":707,"text":708},"11c0035ebd0d",[94],"(2) a per-capability tool allow-list",{"_key":710,"_type":66,"marks":711,"text":712},"55ced279ec63",[]," — it can only call tools on an explicit list scoped to this capability; ",{"_key":714,"_type":66,"marks":715,"text":716},"a99a8ed92329",[94],"(3) a full per-iteration trace",{"_key":718,"_type":66,"marks":719,"text":720},"4e58a13c6c99",[]," — every step recorded, so the decision is replayable; and ",{"_key":722,"_type":66,"marks":723,"text":724},"9f118289d0eb",[94],"(4) the same exits as everything else",{"_key":726,"_type":66,"marks":727,"text":728},"3853860ae120",[]," — the loop’s ",{"_key":730,"_type":66,"marks":731,"text":732},"0f331e54d0db",[72],"output",{"_key":734,"_type":66,"marks":735,"text":736},"ae1c1cb39901",[]," still flows through guardrails, verification, confidence, and routing, and still ",{"_key":738,"_type":66,"marks":739,"text":740},"6bb10e8ce42c",[72],"proposes",{"_key":742,"_type":66,"marks":743,"text":744},"a7385ffc0930",[],", never acts.",[],{"_key":747,"_type":152,"code":748,"language":154,"markDefs":12},"2b0a8c0877f8","class DisallowedTool(Exception): ...    # raised when the loop reaches for an off-list tool\n\nALLOWED_TOOLS = {                       # rail 2: per-capability allow-list, not \"all tools\"\n    \"enrich_request\": {\"search_kb\", \"fetch_record\", \"lookup_reference\"},\n}\n\n@dataclass\nclass Step:                              # rail 3: one trace row per iteration\n    i: int; action: str; args: dict; result_digest: str\n\ndef bounded_react(state, deps, capability, MAX_STEPS=6) -> Proposal:\n    allow = ALLOWED_TOOLS[capability]\n    trace: list[Step] = []\n    for i in range(MAX_STEPS):                                   # rail 1: hard cap\n        action = deps.model.next_action(state, tools=sorted(allow))\n        if action.is_final:                                     # check FIRST — a final answer carries no tool\n            return finalize(action.proposal, trace)             # rail 4: still a PROPOSAL\n        if action.tool not in allow:                            # defense in depth\n            raise DisallowedTool(action.tool)\n        result = deps.tools[action.tool](**action.args)\n        trace.append(Step(i, action.tool, action.args, digest(result)))\n        state = state.with_observation(result)                  # loop-local state type, not the fixed-graph GraphState\n    return escalate(\"hit step cap\", trace)                      # bounded: cap hit → returns a Proposal routed \"hitl_required\"",{"_key":750,"_type":62,"children":751,"markDefs":764,"style":75},"cd1aaed38a37",[752,756,760],{"_key":753,"_type":66,"marks":754,"text":755},"1441e0ee09cd",[],"The ",{"_key":757,"_type":66,"marks":758,"text":759},"8bb5467fb304",[152],"trace",{"_key":761,"_type":66,"marks":762,"text":763},"a42efe6aba82",[]," rides into the ledger entry, so a loop-based decision is exactly as reconstructable as a fixed-graph one. The two rails worth a test each — it always terminates, and it can't reach an off-list tool:",[],{"_key":766,"_type":152,"code":767,"language":154,"markDefs":12},"e3af44428091","def test_always_terminates():\n    deps = fake_deps(model=never_final)                 # a model that never returns is_final\n    p = bounded_react(state, deps, \"enrich_request\", MAX_STEPS=3)\n    assert p.routing == \"hitl_required\"                 # hit the cap → routed to a human, didn't hang\n\ndef test_disallowed_tool_refused():\n    deps = fake_deps(model=calls(\"danger_tool\"))        # a tool not on the allow-list\n    with pytest.raises(DisallowedTool):\n        bounded_react(state, deps, \"enrich_request\")",{"_key":769,"_type":62,"children":770,"markDefs":783,"style":75},"525bd81c7a36",[771,775,779],{"_key":772,"_type":66,"marks":773,"text":774},"0825b3eee388",[],"The decision rule for whether you even need this: ",{"_key":776,"_type":66,"marks":777,"text":778},"f05367ad579b",[72],"does reaching the decision require steps whose number and order depend on what’s discovered along the way?",{"_key":780,"_type":66,"marks":781,"text":782},"e68f6b54346d",[]," No → fixed graph (most capabilities). Yes → bounded loop, four rails, documented as an exception. If you reach for a loop “to be safe” or “for flexibility,” stop — that’s usually the fixed-graph case in a costume. Flexibility you don’t need is nondeterminism you’ll debug at 2am.",[],{"_key":785,"_type":62,"children":786,"markDefs":791,"style":117},"5a25e576a5b9",[787],{"_key":788,"_type":66,"marks":789,"text":790},"f352ce25a9a1",[],"Anti-patterns",[],{"_key":793,"_type":62,"children":794,"level":197,"listItem":198,"markDefs":811,"style":75},"8a33a69e3493",[795,799,803,807],{"_key":796,"_type":66,"marks":797,"text":798},"f85bff2b6483",[94],"The agent writes business state “just this once.”",{"_key":800,"_type":66,"marks":801,"text":802},"ba5b38a03cd5",[]," Now it’s not pure, not testable, and a bug is an incident instead of a bad proposal. Keep the substrate the ",{"_key":804,"_type":66,"marks":805,"text":806},"091ae1df1588",[72],"only",{"_key":808,"_type":66,"marks":809,"text":810},"7e93a2eea99a",[]," mutator.",[],{"_key":813,"_type":62,"children":814,"level":197,"listItem":198,"markDefs":823,"style":75},"b79042799986",[815,819],{"_key":816,"_type":66,"marks":817,"text":818},"499d8a459712",[94],"Implicit state",{"_key":820,"_type":66,"marks":821,"text":822},"a39657794f44",[]," passed as ad-hoc tuples\u002Fdicts between steps — you lose the ability to test a node in isolation. Make the state a typed object.",[],{"_key":825,"_type":62,"children":826,"level":197,"listItem":198,"markDefs":835,"style":75},"c33130a7daad",[827,831],{"_key":828,"_type":66,"marks":829,"text":830},"5b66bfa54e7f",[94],"“Return JSON” in the prompt with no schema enforcement",{"_key":832,"_type":66,"marks":833,"text":834},"7d715e8ac84d",[]," — you’ll still get prose, fences, or trailing commentary. Use schema\u002Ftool\u002Fgrammar enforcement, then validate and retry; constrained ≠ guaranteed.",[],{"_key":837,"_type":62,"children":838,"level":197,"listItem":198,"markDefs":847,"style":75},"b5ec4c417e52",[839,843],{"_key":840,"_type":66,"marks":841,"text":842},"b2d55b56eff6",[94],"“Confidence” lifted from the model.",{"_key":844,"_type":66,"marks":845,"text":846},"5137bdfc5af2",[]," Miscalibrated; compose it from independent signals instead.",[],{"_key":849,"_type":62,"children":850,"level":197,"listItem":198,"markDefs":866,"style":75},"f5f8ac8326db",[851,855,859,862],{"_key":852,"_type":66,"marks":853,"text":854},"b754748f4ac7",[94],"Mutable or skipped audit.",{"_key":856,"_type":66,"marks":857,"text":858},"fbf3313d763e",[]," If you can ",{"_key":860,"_type":66,"marks":861,"text":645},"34048922d9d0",[152],{"_key":863,"_type":66,"marks":864,"text":865},"7631ce183b6e",[]," the ledger it's not an audit trail; if you drop the write under load you have no record of record.",[],{"_key":868,"_type":62,"children":869,"level":197,"listItem":198,"markDefs":886,"style":75},"1a48ff226ef6",[870,874,878,882],{"_key":871,"_type":66,"marks":872,"text":873},"2ed258b160d0",[94],"Unbounded loop “for flexibility.”",{"_key":875,"_type":66,"marks":876,"text":877},"1bfccac05a8e",[]," Default to the fixed graph. When you do loop, use a ",{"_key":879,"_type":66,"marks":880,"text":881},"bc71b49a1174",[152],"for cap and a per-capability allow-list — never ",{"_key":883,"_type":66,"marks":884,"text":885},"95a305e845c7",[],"while not done` or “all tools available.”",[],{"_key":888,"_type":62,"children":889,"markDefs":894,"style":117},"f08678d0bc06",[890],{"_key":891,"_type":66,"marks":892,"text":893},"7c6361f19c1d",[],"The takeaway",[],{"_key":896,"_type":62,"children":897,"markDefs":909,"style":75},"c15df7efd167",[898,902,905],{"_key":899,"_type":66,"marks":900,"text":901},"bbb6d5bf76d9",[],"Level 1 is one discipline applied consistently: contain the model to one node, make the agent a pure function that only ",{"_key":903,"_type":66,"marks":904,"text":740},"98238f99e743",[72],{"_key":906,"_type":66,"marks":907,"text":908},"f979d76334ba",[],", constrain that node to a validated schema, compose confidence from independent signals to route the close calls to a human, let a dumb substrate be the sole mutator after approval, and write every decision to an append-only ledger. When a capability genuinely needs autonomy, budget it — a step cap, a tool allow-list, a trace, the same output checks as everything else.",[],{"_key":911,"_type":62,"children":912,"markDefs":917,"style":75},"a017f59fe152",[913],{"_key":914,"_type":66,"marks":915,"text":916},"4730b9540c42",[],"The model still does what it’s uniquely good at — judgment on messy inputs — but the system around it is deterministic, testable, and auditable. Boring, in the best way: predictable enough to test, contained enough to trust, and defensible enough to ship. That’s the floor everything else in this series is built on.",[],{"_key":919,"_type":62,"children":920,"markDefs":925,"style":75},"5db958bf3f22",[921],{"_key":922,"_type":66,"marks":923,"text":924},"cd15bc91df18",[72],"Series: Running LLM systems in production — Level 1 of 6: Determinism.",[],true,"2026\u002F10\u002F07",{"_type":929,"alt":930,"asset":931},"image","A diagram titled \"Make Your AI Agents Boring,\" detailing a deterministic, linear workflow for AI agents, from entry through reasoning, guardrails, human review, and record-keeping, leading to approval or rejection.",{"_ref":932,"_type":933},"image-a5c8228cbc74e5c74e42a697b184129575a925a1-3000x1500-png","reference","2026-10-07T20:14:49.390Z",{"_type":936,"canonicalUrl":937},"seo","https:\u002F\u002Fmedium.com\u002F@varunjindal9\u002Fmake-your-ai-agents-boring-the-determinism-layer-bae580d88c89",{"_type":10,"current":939},"part-1-make-your-ai-agents-boring-the-determinism-layer",[941,950,961,984],{"_createdAt":942,"_id":943,"_rev":944,"_type":945,"_updatedAt":946,"slug":947,"title":949},"2023-05-23T16:43:21Z","wp-tagcat-ai","fpDTFQqIDjNJIbHDKPBGpV","blogTag","2025-01-30T16:19:01Z",{"current":948},"ai","AI",{"_createdAt":951,"_id":952,"_rev":953,"_system":954,"_type":945,"_updatedAt":957,"slug":958,"title":960},"2026-06-12T16:16:20Z","51c761d7-73f7-42f4-aa49-8484e3849e7c","P0qLqkXH0zpkT6RRZ9Iwel",{"base":955},{"id":952,"rev":956},"MwgZb85ftkde1TTvQsHYa6","2026-09-28T16:40:45Z",{"_type":10,"current":959},"building-software","Building software",{"_createdAt":962,"_id":963,"_rev":964,"_system":965,"_type":945,"_updatedAt":968,"description":969,"featuredPosts":978,"slug":981,"title":983},"2025-04-24T16:28:57Z","797b8797-6e65-4723-b53f-8bc005305384","46s78gX2DRxVswzX0kQ1Ty",{"base":966},{"id":963,"rev":967},"IpfPEqg1c3Byvj9RrB3Xaj","2026-10-07T20:13:00Z",[970],{"_key":971,"_type":62,"children":972,"markDefs":977,"style":75},"bb32f75814b4",[973],{"_key":974,"_type":66,"marks":975,"text":976},"dbcf27ef29b3",[],"Community-generated articles submitted for your reading pleasure. If you’re interested in seeing your work here, log in with your Stack Overflow account and click the link below. Articles will be licensed under a CC BY-SA 4.0 grant. ",[],[979],{"_key":980,"_type":933},"9d9ea8c4082d",{"_type":10,"current":982},"contributed","The Heap",{"_createdAt":985,"_id":986,"_rev":987,"_system":988,"_type":945,"_updatedAt":991,"description":992,"slug":1012,"title":1014},"2025-08-08T15:49:22Z","39391cf4-6f9a-4238-8670-c1e44b66db9e","09X6HDzCi2VfMov6gSLf7H",{"base":989},{"id":986,"rev":990},"TdCcmC7LyfLVwjB8GEXoh6","2025-12-10T19:34:33Z",[993,1001],{"_key":994,"_type":62,"children":995,"markDefs":1000,"style":75},"a4b1a37cbbcc",[996],{"_key":997,"_type":66,"marks":998,"text":999},"d8e8f3e0fd9c",[],"These articles are licensed under a Creative Commons Attribution-ShareAlike 4.0 International license. ",[],{"_key":1002,"_type":62,"children":1003,"markDefs":1009,"style":75},"7effd489c71f",[1004],{"_key":1005,"_type":66,"marks":1006,"text":1008},"538808bb5325",[1007],"fd643b288690","creativecommons.org\u002Flicenses\u002Fby-sa\u002F4.0\u002Fdeed.en",[1010],{"_key":1007,"_type":1011},"link",{"_type":10,"current":1013},"cc-by-sa","CC BY-SA 4.0","Part 1: Make your AI agents boring: the determinism layer",[1017,1023,1029,1035],{"_id":1018,"publishedAt":1019,"slug":1020,"sponsored":12,"title":1022},"ce1fd642-fe2b-4d83-95ad-67b7645d7959","2026-10-07T20:40:29.066Z",{"_type":10,"current":1021},"part-3-knowing-when-your-agent-doesn-t-know-the-confidence-layer","Part 3: Knowing when your agent doesn’t know: the confidence layer",{"_id":1024,"publishedAt":1025,"slug":1026,"sponsored":12,"title":1028},"f17b27e9-17f5-4203-8066-68c8df26ef42","2026-10-07T20:31:12.368Z",{"_type":10,"current":1027},"evals-as-a-deployment-gate-and-how-to-know-when-they-drift","Part 2: Evals as a deployment gate — and how to know when they drift",{"_id":1030,"publishedAt":1031,"slug":1032,"sponsored":12,"title":1034},"f2b2b7a0-c8e2-4387-8a7d-439b28bab359","2026-10-07T20:03:47.355Z",{"_type":10,"current":1033},"implementing-a-modular-master-agent-telemetry-and-diagnostic-framework-in-python-prime-sentinel-command-psc","Implementing a Modular Master-Agent Telemetry & Diagnostic Framework in Python: Prime-Sentinel Command (PSC)",{"_id":1036,"publishedAt":1037,"slug":1038,"sponsored":12,"title":1040},"b1af6214-1a59-4309-8484-920746d31d04","2026-10-06T14:00:00.000Z",{"_type":10,"current":1039},"the-results-of-the-2026-developer-survey-are-here","The results of the 2026 Developer Survey are here!",{"data":1042,"sourceMap":-1},{"count":1043,"lastTimestamp":12},0]