[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"sanity-H_1tqIrpqGfgJTvD362-8a8UIWjWrhtU2X35KN3Hvuw":3,"sanity-TYZFQN6xOwUzJP7buxvG756lzyTIgsD0eqZoDS67_1Y":729},{"data":4,"sourceMap":-1},{"latestPodcast":5,"latestReleases":14,"post":39,"recent":704},[6],{"_id":7,"publishedAt":8,"slug":9,"sponsored":12,"title":13},"8e6c3a8a-3d27-44c8-be90-0b74d1090a4a","2026-10-06T17:00:00.000Z",{"_type":10,"current":11},"slug","tales-from-the-2026-developer-survey-results",null,"Tales from the 2026 Developer Survey results",[15,21,27,33],{"_id":16,"publishedAt":17,"slug":18,"title":20},"7a2d88e5-ee53-4f7c-a46e-47bed1cebadf","2026-09-30T16:00:00.000Z",{"_type":10,"current":19},"anyone-can-start-building-verified-knowledge-with-stack-internal","Anyone can start building verified knowledge with Stack Internal",{"_id":22,"publishedAt":23,"slug":24,"title":26},"12c6a8a7-135f-401f-ac1a-1c26c33e69c0","2026-09-03T16:00:00.000Z",{"_type":10,"current":25},"security-control-and-accessibility-si-2026-6","Elevating security, control, and accessibility: Stack Internal 2026.6",{"_id":28,"publishedAt":29,"slug":30,"title":32},"adcf1bca-3295-4ac5-9b3c-23f337974190","2026-07-30T15:10:00.000Z",{"_type":10,"current":31},"introducing-stack-internal-new-platform-experience","Your trusted knowledge layer: Introducing Stack Internal's new platform experience",{"_id":34,"publishedAt":35,"slug":36,"title":38},"eb5b66eb-9410-4329-83bb-22bbff39402a","2026-04-28T13:00:00.000Z",{"_type":10,"current":37},"turn-scattered-knowledge-into-trusted-intelligence","Turning scattered knowledge into trusted intelligence: Stack Internal 2026.3",{"_createdAt":40,"_id":41,"_rev":42,"_system":43,"_type":46,"_updatedAt":47,"author":48,"body":59,"comments":614,"dateUrl":615,"image":616,"product":12,"publishedAt":622,"seo":623,"slug":626,"sponsored":12,"tags":628,"title":703,"visible":614},"2026-10-07T20:31:12Z","f17b27e9-17f5-4203-8066-68c8df26ef42","46s78gX2DRxVswzX0l5qdy",{"base":44},{"id":41,"rev":45},"46s78gX2DRxVswzX0l1Ypy","blogPost","2026-10-07T20:50:06Z",[49],{"_createdAt":50,"_id":51,"_rev":52,"_type":53,"_updatedAt":54,"employee":55,"name":56,"slug":57},"2026-10-07T20:10:11Z","c7434ebf-211a-49d5-a564-22a91d63ed1e","9fZfA2bz90M2vN1VVZP5Ov","blogAuthor","2026-10-07T20:12:01Z","none","Varun Jindal",{"_type":10,"current":58},"varun-jindal",[60,76,93,101,110,115,131,139,147,151,159,167,171,174,182,190,214,222,230,244,256,264,272,288,291,294,325,333,336,352,360,376,379,387,395,398,406,419,431,443,463,471,483,502,514,526,538,550,562,574,582,606],{"_key":61,"_type":62,"children":63,"markDefs":74,"style":75},"e9e77c73209f","block",[64,69],{"_key":65,"_type":66,"marks":67,"text":68},"c95ed37eb921","span",[],"If you can deploy a prompt change without an eval failing the build, you don’t have evals — you have a notebook. And once the gate is green, the slow leaks are still coming for you. Here’s the gate, the baseline, and the shadow-",{"_key":70,"_type":66,"marks":71,"text":73},"79316d9f2507",[72],"em","eval loop that catch both.",[],"normal",{"_key":77,"_type":62,"children":78,"markDefs":92,"style":75},"56d1dc84755d",[79,83,88],{"_key":80,"_type":66,"marks":81,"text":82},"9023f30837c5",[],"This is Level 2 of the maturity model: ",{"_key":84,"_type":66,"marks":85,"text":87},"8451f84591be",[86],"strong","evaluation",{"_key":89,"_type":66,"marks":90,"text":91},"ba32cd553a08",[],". The principle is short — you don’t ship on hope, you ship on a gate, and then you watch for drift afterward. A gate protects the moment of deploy. Drift detection protects the weeks in between. You need both, and they’re built from different machinery.",[],{"_key":94,"_type":62,"children":95,"markDefs":100,"style":75},"08d864ba15b8",[96],{"_key":97,"_type":66,"marks":98,"text":99},"38cda2c6b783",[],"Most teams “do evals” the way they once “did tests” before CI: occasionally, manually, in a notebook, after something already went wrong. For LLM systems that’s not enough, because the thing most likely to silently break your product isn’t a code change — it’s a prompt tweak, a model upgrade, or a temperature nudge that looks harmless and quietly tanks quality on 8% of inputs. The fix is to treat evaluation like tests: a gate that runs in CI and fails the build on regression.",[],{"_key":102,"_type":62,"children":103,"markDefs":108,"style":109},"c1261e8af64e",[104],{"_key":105,"_type":66,"marks":106,"text":107},"e562a276154b",[],"A golden case is data",[],"h3",{"_key":111,"_type":112,"code":113,"language":114,"markDefs":12},"8878ec1d5502","code","{\n  \"id\": \"case-0412\",\n  \"slice\": \"tier1\u002Fen\",\n  \"input\": { \"request\": \"...\", \"context\": { } },\n  \"expect\": { \"decision\": \"approve\", \"min_confidence\": 0.80 },\n  \"must_not\": { \"decision\": \"auto_resolve\" },\n  \"tags\": [\"edge-case\", \"previously-broke\"]\n}","json",{"_key":116,"_type":62,"children":117,"markDefs":130,"style":75},"aa5d694e2eb0",[118,122,126],{"_key":119,"_type":66,"marks":120,"text":121},"0086707b83fd",[],"Curate it like tests: cover the common cases, the known-hard cases, and ",{"_key":123,"_type":66,"marks":124,"text":125},"489f99c5a3c5",[86],"every past failure you’ve fixed",{"_key":127,"_type":66,"marks":128,"text":129},"87605e17d506",[]," (regression cases never get deleted). It lives in version control next to the code, so a change to the set is a reviewable diff.",[],{"_key":132,"_type":62,"children":133,"markDefs":138,"style":109},"f627d8058cc4",[134],{"_key":135,"_type":66,"marks":136,"text":137},"d0571a23f270",[],"The scorer",[],{"_key":140,"_type":62,"children":141,"markDefs":146,"style":75},"5390816156c2",[142],{"_key":143,"_type":66,"marks":144,"text":145},"1c9668fc0c49",[],"Per case, run the current prompt\u002Fmodel and score the result. Keep scoring explicit and typed — a mix of exact assertions and tolerances.",[],{"_key":148,"_type":112,"code":149,"language":150,"markDefs":12},"6c759cdca748","from dataclasses import dataclass\n\n@dataclass\nclass CaseResult:\n    id: str; slice: str; passed: bool; checks: list\n\ndef score_case(case, output) -> CaseResult:\n    exp = case.get(\"expect\", {})\n    checks = []\n    if \"decision\" in exp:                                                       # symmetric .get on both sides\n        checks.append((\"decision\", output.get(\"decision\") == exp[\"decision\"]))\n    if \"min_confidence\" in exp:\n        checks.append((\"min_conf\", output.get(\"confidence\", 0) >= exp[\"min_confidence\"]))\n    for k, v in case.get(\"must_not\", {}).items():\n        checks.append((f\"not_{k}\", output.get(k) != v))\n    assert checks, \"empty expect + no must_not would vacuously pass\"            # now this CAN actually fire\n    return CaseResult(case[\"id\"], case[\"slice\"], all(ok for _, ok in checks), checks)\n\ndef test_scorer_catches_regression():                  # sensitivity: it MUST fail a wrong decision\n    case = {\"id\": \"c1\", \"slice\": \"tier1\u002Fen\", \"expect\": {\"decision\": \"approve\", \"min_confidence\": 0.8}}\n    assert     score_case(case, {\"decision\": \"approve\", \"confidence\": 0.9}).passed\n    assert not score_case(case, {\"decision\": \"reject\",  \"confidence\": 0.9}).passed","python",{"_key":152,"_type":62,"children":153,"markDefs":158,"style":109},"bff1dfe7931e",[154],{"_key":155,"_type":66,"marks":156,"text":157},"bdd385ee2301",[],"The gate",[],{"_key":160,"_type":62,"children":161,"markDefs":166,"style":75},"bef814fc67ba",[162],{"_key":163,"_type":66,"marks":164,"text":165},"f497bb8d47fd",[],"One command, wired into CI, that fails the build when the pass rate (or any keyed metric) drops below the bar.",[],{"_key":168,"_type":112,"code":169,"language":170,"markDefs":12},"cf7685c03fbd","# .ci\u002Feval-gate.yml  (illustrative)\neval-gate:\n  steps:\n    - run: eval run --suite golden --gate --min-pass 0.98 --baseline artifacts\u002Fbaseline.json\n  # exit code != 0 fails the pipeline → deploy blocked","yaml",{"_key":172,"_type":112,"code":173,"language":170,"markDefs":12},"93cce859b669","$ eval run --suite golden --gate --min-pass 0.98\n  cases: 240   passed: 236   failed: 4   pass-rate: 98.3%\n  gate(min-pass>=0.98): PASS\n  failures: case-0412 (decision), case-0511 (min_conf), ...",{"_key":175,"_type":62,"children":176,"markDefs":181,"style":75},"a9ab971bab00",[177],{"_key":178,"_type":66,"marks":179,"text":180},"65e7b4adb239",[],"Now a prompt change that improves one scenario but breaks two others can’t ship silently — the author sees it in their PR, not a customer three weeks later.",[],{"_key":183,"_type":62,"children":184,"markDefs":189,"style":109},"35be8321205d",[185],{"_key":186,"_type":66,"marks":187,"text":188},"8cbc29da1b87",[],"Route cases to the right slice automatically",[],{"_key":191,"_type":62,"children":192,"markDefs":213,"style":75},"9f67aa07851a",[193,197,201,205,209],{"_key":194,"_type":66,"marks":195,"text":196},"4af907723737",[],"As the system grows, “the golden set” becomes many sub-suites for different request types, segments, or locales. Don’t make a human pick which suite to run — tag each case with its ",{"_key":198,"_type":66,"marks":199,"text":200},"84380d9468c0",[112],"slice",{"_key":202,"_type":66,"marks":203,"text":204},"9bca69569b10",[]," and let the runner ",{"_key":206,"_type":66,"marks":207,"text":208},"c1227b9a7daa",[86],"route cases to the matching slice using the same attributes the system uses at runtime",{"_key":210,"_type":66,"marks":211,"text":212},"d8a421a5af40",[],". The eval for a slice runs the cases that belong to it; no manual selection, no drift between \"what we test\" and \"what we run.\" And don't gate on one global pass rate: a 99% overall can hide a 70% slice, so gate per slice too.",[],{"_key":215,"_type":62,"children":216,"markDefs":221,"style":109},"ae9cac1bb63b",[217],{"_key":218,"_type":66,"marks":219,"text":220},"74b89f45fa45",[],"Two kinds of regression",[],{"_key":223,"_type":62,"children":224,"markDefs":229,"style":75},"120b9ae2a83c",[225],{"_key":226,"_type":66,"marks":227,"text":228},"934ab31fb1b8",[],"The gate above is excellent at one thing and blind to another:",[],{"_key":231,"_type":62,"children":232,"level":241,"listItem":242,"markDefs":243,"style":75},"4778b2e93337",[233,237],{"_key":234,"_type":66,"marks":235,"text":236},"b3330c58f6f2",[86],"Cliffs",{"_key":238,"_type":66,"marks":239,"text":240},"6f61397b6984",[]," — a change drops quality sharply and immediately. The pre-deploy gate catches these.",1,"bullet",[],{"_key":245,"_type":62,"children":246,"level":241,"listItem":242,"markDefs":255,"style":75},"87f7741e5dbe",[247,251],{"_key":248,"_type":66,"marks":249,"text":250},"53b080cefd98",[86],"Slow leaks",{"_key":252,"_type":66,"marks":253,"text":254},"9c5b116c233a",[]," — 0.995 → 0.99 → 0.985 over three releases, each step too small to trip the 0.98 floor. The gate waves each one through; by the time it’s obviously bad you’ve shipped it five times.",[],{"_key":257,"_type":62,"children":258,"markDefs":263,"style":75},"8c56cb264ea5",[259],{"_key":260,"_type":66,"marks":261,"text":262},"4515bbd1a944",[],"You can have a green eval gate and a system that’s quietly getting worse. Gates test the inputs you thought of, at the moment you deploy. Production changes underneath you: inputs shift, the provider updates a model, a prompt tweak helps the cases you tested and hurts the ones you didn’t. Drift detection defends against the second kind. It needs different machinery.",[],{"_key":265,"_type":62,"children":266,"markDefs":271,"style":109},"0b5c24a55520",[267],{"_key":268,"_type":66,"marks":269,"text":270},"ee16a9fe8df3",[],"Baselines: compare to known-good, not just a floor",[],{"_key":273,"_type":62,"children":274,"markDefs":287,"style":75},"2862fa3e3f12",[275,279,283],{"_key":276,"_type":66,"marks":277,"text":278},"ed9e658133c0",[],"A gate asks “above the line?” Drift asks “worse than before?” So capture a ",{"_key":280,"_type":66,"marks":281,"text":282},"070ea5f7d910",[86],"baseline",{"_key":284,"_type":66,"marks":285,"text":286},"f6a420d27a3d",[]," — the metrics from the last known-good release — and compare every run to it.",[],{"_key":289,"_type":112,"code":290,"language":114,"markDefs":12},"ffe5ae8f8a7c","\u002F\u002F baseline.json — captured deliberately when you bless a release as known-good\n{\n  \"release\": \"2026-01-09\",\n  \"metrics\": {\n    \"exact_match\":        { \"mean\": 0.942, \"n\": 2400, \"std\": 0.012 },\n    \"mean_confidence\":    { \"mean\": 0.871, \"n\": 2400, \"std\": 0.030 },\n    \"human_override_rate\":{ \"mean\": 0.060, \"n\": 2400 }\n  }\n}",{"_key":292,"_type":112,"code":293,"language":150,"markDefs":12},"433aca22a462","import math\n# per-metric minimum meaningful change — a 0.02 move matters for a 0.06 rate, trivial for confidence\nMIN_EFFECT = {\"exact_match\": 0.01, \"human_override_rate\": 0.02, \"mean_confidence\": 0.03}\nHIGHER_IS_WORSE = {\"human_override_rate\", \"judge_disagreement_ratio\", \"abstention_rate\"}  # for these, UP = worse\n\ndef compare_to_baseline(metric, cur, base, is_proportion=True):\n    # cur\u002Fbase are {\"mean\":.., \"n\":.., \"std\":..}. Use BOTH runs' n — not just the baseline's.\n    delta = cur[\"mean\"] - base[\"mean\"]\n    if is_proportion:                                   # rate metrics → two-proportion z-test\n        p  = (base[\"mean\"]*base[\"n\"] + cur[\"mean\"]*cur[\"n\"]) \u002F (base[\"n\"] + cur[\"n\"])   # pooled\n        se = math.sqrt(p*(1-p) * (1\u002Fbase[\"n\"] + 1\u002Fcur[\"n\"]))\n    else:                                               # continuous → SE of a DIFFERENCE of means\n        se = math.sqrt(base[\"std\"]**2\u002Fbase[\"n\"] + cur[\"std\"]**2\u002Fcur[\"n\"])\n    z = delta\u002Fse if se else 0.0\n    worse = delta > 0 if metric in HIGHER_IS_WORSE else delta \u003C 0   # direction is per-metric, not always \"lower\"\n    drifted = worse and abs(delta) >= MIN_EFFECT.get(metric, 0.02) and abs(z) > 2   # meaningful + ~2σ the bad way\n    return {\"metric\": metric, \"delta\": round(delta, 4), \"z\": round(z, 1), \"drift\": drifted}",{"_key":295,"_type":62,"children":296,"markDefs":324,"style":75},"157c4b5dac2a",[297,301,304,308,312,316,320],{"_key":298,"_type":66,"marks":299,"text":300},"d5d4e23df289",[],"The common bug: dividing the ",{"_key":302,"_type":66,"marks":303,"text":282},"b3b14cc3ad04",[72],{"_key":305,"_type":66,"marks":306,"text":307},"1dbca45eac0d",[]," std by ",{"_key":309,"_type":66,"marks":310,"text":311},"aa74fe277249",[112],"sqrt(n)",{"_key":313,"_type":66,"marks":314,"text":315},"2aec4692db74",[]," ignores the current run's own size and variance — sample 50 live decisions against a 2,400-row baseline and you'll flag pure noise. The SE of a difference uses both; proportions get the pooled two-proportion form, not a stored ",{"_key":317,"_type":66,"marks":318,"text":319},"9332a71639ba",[112],"std",{"_key":321,"_type":66,"marks":322,"text":323},"8a359cb2d84c",[],".",[],{"_key":326,"_type":62,"children":327,"markDefs":332,"style":75},"e42cd23c7162",[328],{"_key":329,"_type":66,"marks":330,"text":331},"78e3e2dbbaf0",[],"A change can be above your floor and still meaningfully below baseline — that’s the signal the gate misses:",[],{"_key":334,"_type":112,"code":335,"markDefs":12},"627e6a20d883","metric         baseline  current  delta\nexact-match    0.942     0.913   -0.029  ⚠ regression vs baseline (still > floor, but flagged)",{"_key":337,"_type":62,"children":338,"markDefs":351,"style":75},"80574d12f95f",[339,343,347],{"_key":340,"_type":66,"marks":341,"text":342},"3b13422ddad9",[],"Re-baseline ",{"_key":344,"_type":66,"marks":345,"text":346},"685e70cbdff9",[72],"deliberately",{"_key":348,"_type":66,"marks":349,"text":350},"0666235b4649",[]," (when you validate a new known-good), never automatically, or you let drift become the new normal.",[],{"_key":353,"_type":62,"children":354,"markDefs":359,"style":109},"dbb433268451",[355],{"_key":356,"_type":66,"marks":357,"text":358},"008ce8cb3bdd",[],"Shadow evals: grade live traffic",[],{"_key":361,"_type":62,"children":362,"markDefs":375,"style":75},"d186082de754",[363,367,371],{"_key":364,"_type":66,"marks":365,"text":366},"39e141d91a6e",[],"Your golden set is finite and curated; production is infinite and surprising. A ",{"_key":368,"_type":66,"marks":369,"text":370},"eb78175e5647",[86],"shadow eval",{"_key":372,"_type":66,"marks":373,"text":374},"e7249915251a",[]," samples real (de-identified) inputs, scores the actual decisions, and tracks the pass rate over time.",[],{"_key":377,"_type":112,"code":378,"language":150,"markDefs":12},"95c5a71a6629","SHADOW_SAMPLE_RATE = 0.05    # score 5% of live decisions out-of-band (config, not hardcoded)\n\ndef maybe_shadow(decision, sample_rate=SHADOW_SAMPLE_RATE):\n    # deterministic_hash = your stable id hash; score_against = your rules\u002Fjudge scoring entrypoint\n    # % 10000 (not % 100) so sub-1% rates don't silently round to zero sampling\n    if deterministic_hash(decision[\"decision_id\"]) % 10000 \u003C sample_rate * 10000:\n        verdict = score_against(rules_or_judge, decision)   # async, off the hot path\n        emit(\"shadow_eval_pass_ratio\", 1.0 if verdict.ok else 0.0, slice=decision[\"slice\"])\n        if not verdict.ok:\n            add_to_review_queue(decision)                   # candidate new golden case",{"_key":380,"_type":62,"children":381,"markDefs":386,"style":75},"ae02b65dbb66",[382],{"_key":383,"_type":66,"marks":384,"text":385},"22d7ae07dd97",[],"Failures become new golden cases — production hardens your suite exactly where reality is hardest. Use a deterministic hash of the id, not RNG, so sampling is reproducible and doesn’t make tests flaky.",[],{"_key":388,"_type":62,"children":389,"markDefs":394,"style":109},"9873c510c87c",[390],{"_key":391,"_type":66,"marks":392,"text":393},"57b5b2a9461d",[],"Early-warning signals (watch the trend, not the point)",[],{"_key":396,"_type":112,"code":397,"markDefs":12},"9c1741690236","signal                         drift smell\n-----------------------------  ----------------------------------------------------------------------------\n`human_override_rate` ↑        humans reversing the agent more — quality slips before metrics fully show it\nconfidence distribution shift  trending down (less sure), or up while accuracy falls (miscalibration)\n`judge_disagreement_ratio` ↑   the second model overruling the first more often\nabstention rate ↑              the system punting more than it used to",{"_key":399,"_type":62,"children":400,"markDefs":405,"style":109},"08c986b40c71",[401],{"_key":402,"_type":66,"marks":403,"text":404},"7449f0fe5812",[],"When it fires — investigate, don’t auto-rollback",[],{"_key":407,"_type":62,"children":408,"level":241,"listItem":417,"markDefs":418,"style":75},"c714f4b20a0b",[409,413],{"_key":410,"_type":66,"marks":411,"text":412},"eef5a2139455",[86],"Confirm",{"_key":414,"_type":66,"marks":415,"text":416},"a5bd6a3cafb7",[]," it’s real (meaningful + ~2σ, not noise).","number",[],{"_key":420,"_type":62,"children":421,"level":241,"listItem":417,"markDefs":430,"style":75},"782e88492d6d",[422,426],{"_key":423,"_type":66,"marks":424,"text":425},"445fa59131a5",[86],"Localize",{"_key":427,"_type":66,"marks":428,"text":429},"64a62a14b63a",[]," it (which slice\u002Fcapability — your tagged evals + per-slice metrics tell you).",[],{"_key":432,"_type":62,"children":433,"level":241,"listItem":417,"markDefs":442,"style":75},"ae0726409abc",[434,438],{"_key":435,"_type":66,"marks":436,"text":437},"09ba58f3ea0b",[86],"Find the cause",{"_key":439,"_type":66,"marks":440,"text":441},"75a3314d8ab5",[]," (model update? prompt change? input shift?).",[],{"_key":444,"_type":62,"children":445,"level":241,"listItem":417,"markDefs":462,"style":75},"e5f8142956d7",[446,450,454,458],{"_key":447,"_type":66,"marks":448,"text":449},"e020610babce",[86],"Fix",{"_key":451,"_type":66,"marks":452,"text":453},"a7e70402f0c9",[]," via the normal change process, then ",{"_key":455,"_type":66,"marks":456,"text":457},"5edef50739ef",[86],"re-baseline",{"_key":459,"_type":66,"marks":460,"text":461},"0b7255d3a6e4",[]," once validated.",[],{"_key":464,"_type":62,"children":465,"markDefs":470,"style":109},"de082b117a9f",[466],{"_key":467,"_type":66,"marks":468,"text":469},"58eee7530009",[],"Anti-patterns",[],{"_key":472,"_type":62,"children":473,"level":241,"listItem":242,"markDefs":482,"style":75},"faf37cd9ddbf",[474,478],{"_key":475,"_type":66,"marks":476,"text":477},"a6a5ec776f27",[86],"Evals in a notebook",{"_key":479,"_type":66,"marks":480,"text":481},"afa54dfb334e",[]," — if it’s not in CI failing builds, it won’t run when it matters.",[],{"_key":484,"_type":62,"children":485,"level":241,"listItem":242,"markDefs":501,"style":75},"f0068c984ee2",[486,490,494,498],{"_key":487,"_type":66,"marks":488,"text":489},"ccbfd2c3f18c",[86],"Asserting on the model’s confidence instead of correctness",{"_key":491,"_type":66,"marks":492,"text":493},"d9b849ecb9f9",[]," — measure whether it was ",{"_key":495,"_type":66,"marks":496,"text":497},"ede50c612aa2",[72],"right",{"_key":499,"_type":66,"marks":500,"text":323},"8f41dfed77f8",[],[],{"_key":503,"_type":62,"children":504,"level":241,"listItem":242,"markDefs":513,"style":75},"8b372e99ac3b",[505,509],{"_key":506,"_type":66,"marks":507,"text":508},"5c6c3d113eef",[86],"Deleting fixed-bug cases",{"_key":510,"_type":66,"marks":511,"text":512},"4efe7f16ba1d",[]," — they’re your regression suite; keep them forever.",[],{"_key":515,"_type":62,"children":516,"level":241,"listItem":242,"markDefs":525,"style":75},"0691fdf1a72b",[517,521],{"_key":518,"_type":66,"marks":519,"text":520},"a1d70f84125d",[86],"One global pass rate",{"_key":522,"_type":66,"marks":523,"text":524},"a8644ae2be5c",[]," — a 99% overall can hide a 70% slice; gate per slice too.",[],{"_key":527,"_type":62,"children":528,"level":241,"listItem":242,"markDefs":537,"style":75},"b36ec1c4a239",[529,533],{"_key":530,"_type":66,"marks":531,"text":532},"f4f78e55ab0a",[86],"Floor-only gating",{"_key":534,"_type":66,"marks":535,"text":536},"2b15c4188b78",[]," — you’ll ship the slow leak; baseline against known-good too.",[],{"_key":539,"_type":62,"children":540,"level":241,"listItem":242,"markDefs":549,"style":75},"c171847ccdf4",[541,545],{"_key":542,"_type":66,"marks":543,"text":544},"22db3626a833",[86],"RNG sampling in eval\u002Fjudge paths",{"_key":546,"_type":66,"marks":547,"text":548},"34343dd35aa2",[]," — makes tests flaky; hash the id.",[],{"_key":551,"_type":62,"children":552,"level":241,"listItem":242,"markDefs":561,"style":75},"d0d83e5d2321",[553,557],{"_key":554,"_type":66,"marks":555,"text":556},"adde3153cb0c",[86],"Auto re-baselining",{"_key":558,"_type":66,"marks":559,"text":560},"dd015de0b54d",[]," — silently launders drift into the new normal.",[],{"_key":563,"_type":62,"children":564,"level":241,"listItem":242,"markDefs":573,"style":75},"8bbc8778ac1f",[565,569],{"_key":566,"_type":66,"marks":567,"text":568},"c5ae43986972",[86],"Watching points, not trends",{"_key":570,"_type":66,"marks":571,"text":572},"e27510b681dd",[]," — a single bad hour is noise; a two-week slide is drift.",[],{"_key":575,"_type":62,"children":576,"markDefs":581,"style":109},"f5bf5e9b7695",[577],{"_key":578,"_type":66,"marks":579,"text":580},"f4dc0d678b5b",[],"The takeaway",[],{"_key":583,"_type":62,"children":584,"markDefs":605,"style":75},"d667e9dc3c30",[585,589,593,597,601],{"_key":586,"_type":66,"marks":587,"text":588},"ae571a4359fd",[],"Make evals a gate, not a notebook: cases as versioned data, an explicit scorer, a CI command that fails the build on regression, auto-routing to slices. Then add the part the gate can’t do — baseline every run against known-good with a real significance check, shadow-eval a sample of live traffic into a pass ratio, watch override\u002Fconfidence\u002Fdisagreement trends, and feed production’s surprises back into the golden set. Cliffs are easy; the slow leaks sink quality, and they only show up if you’re watching for ",{"_key":590,"_type":66,"marks":591,"text":592},"eb7a3f03e92e",[72],"worse than before",{"_key":594,"_type":66,"marks":595,"text":596},"80b0017cdc22",[],", not just ",{"_key":598,"_type":66,"marks":599,"text":600},"0df6e0bd1257",[72],"below the line",{"_key":602,"_type":66,"marks":603,"text":604},"6592d2160adb",[],". Do both and “ship a prompt change” stops being a gamble and becomes a green check — the only way to move fast on an LLM system without breaking it quietly.",[],{"_key":607,"_type":62,"children":608,"markDefs":613,"style":75},"ada93a9540ed",[609],{"_key":610,"_type":66,"marks":611,"text":612},"8bf796aa6334",[72],"Series: Running LLM systems in production — Level 2 of 6: Evaluation.",[],true,"2026\u002F10\u002F07",{"_type":617,"alt":618,"asset":619},"image","Diagram illustrating \"EVALS AS A DEPLOYMENT GATE\" with a quality control gate, a strip-chart showing quality drifting downward over releases, and inspection notes.",{"_ref":620,"_type":621},"image-1d3b4a8953f24499f8d4efeb586549a0baa73f73-3000x1500-png","reference","2026-10-07T20:31:12.368Z",{"_type":624,"canonicalUrl":625},"seo","https:\u002F\u002Fmedium.com\u002F@varunjindal9\u002Fevals-as-a-deployment-gate-and-how-to-know-when-they-drift-9e4dcc77ef26",{"_type":10,"current":627},"evals-as-a-deployment-gate-and-how-to-know-when-they-drift",[629,638,649,672],{"_createdAt":630,"_id":631,"_rev":632,"_type":633,"_updatedAt":634,"slug":635,"title":637},"2023-05-23T16:43:21Z","wp-tagcat-ai","fpDTFQqIDjNJIbHDKPBGpV","blogTag","2025-01-30T16:19:01Z",{"current":636},"ai","AI",{"_createdAt":639,"_id":640,"_rev":641,"_system":642,"_type":633,"_updatedAt":645,"slug":646,"title":648},"2026-06-12T16:16:20Z","51c761d7-73f7-42f4-aa49-8484e3849e7c","P0qLqkXH0zpkT6RRZ9Iwel",{"base":643},{"id":640,"rev":644},"MwgZb85ftkde1TTvQsHYa6","2026-09-28T16:40:45Z",{"_type":10,"current":647},"building-software","Building software",{"_createdAt":650,"_id":651,"_rev":652,"_system":653,"_type":633,"_updatedAt":656,"description":657,"featuredPosts":666,"slug":669,"title":671},"2025-04-24T16:28:57Z","797b8797-6e65-4723-b53f-8bc005305384","46s78gX2DRxVswzX0kQ1Ty",{"base":654},{"id":651,"rev":655},"IpfPEqg1c3Byvj9RrB3Xaj","2026-10-07T20:13:00Z",[658],{"_key":659,"_type":62,"children":660,"markDefs":665,"style":75},"bb32f75814b4",[661],{"_key":662,"_type":66,"marks":663,"text":664},"dbcf27ef29b3",[],"Community-generated articles submitted for your reading pleasure. If you’re interested in seeing your work here, log in with your Stack Overflow account and click the link below. Articles will be licensed under a CC BY-SA 4.0 grant. ",[],[667],{"_key":668,"_type":621},"9d9ea8c4082d",{"_type":10,"current":670},"contributed","The Heap",{"_createdAt":673,"_id":674,"_rev":675,"_system":676,"_type":633,"_updatedAt":679,"description":680,"slug":700,"title":702},"2025-08-08T15:49:22Z","39391cf4-6f9a-4238-8670-c1e44b66db9e","09X6HDzCi2VfMov6gSLf7H",{"base":677},{"id":674,"rev":678},"TdCcmC7LyfLVwjB8GEXoh6","2025-12-10T19:34:33Z",[681,689],{"_key":682,"_type":62,"children":683,"markDefs":688,"style":75},"a4b1a37cbbcc",[684],{"_key":685,"_type":66,"marks":686,"text":687},"d8e8f3e0fd9c",[],"These articles are licensed under a Creative Commons Attribution-ShareAlike 4.0 International license. ",[],{"_key":690,"_type":62,"children":691,"markDefs":697,"style":75},"7effd489c71f",[692],{"_key":693,"_type":66,"marks":694,"text":696},"538808bb5325",[695],"fd643b288690","creativecommons.org\u002Flicenses\u002Fby-sa\u002F4.0\u002Fdeed.en",[698],{"_key":695,"_type":699},"link",{"_type":10,"current":701},"cc-by-sa","CC BY-SA 4.0","Part 2: Evals as a deployment gate — and how to know when they drift",[705,711,717,723],{"_id":706,"publishedAt":707,"slug":708,"sponsored":12,"title":710},"ce1fd642-fe2b-4d83-95ad-67b7645d7959","2026-10-07T20:40:29.066Z",{"_type":10,"current":709},"part-3-knowing-when-your-agent-doesn-t-know-the-confidence-layer","Part 3: Knowing when your agent doesn’t know: the confidence layer",{"_id":712,"publishedAt":713,"slug":714,"sponsored":12,"title":716},"39f2fc9f-742a-4273-a4d1-abdecd14fcf8","2026-10-07T20:14:49.390Z",{"_type":10,"current":715},"part-1-make-your-ai-agents-boring-the-determinism-layer","Part 1: Make your AI agents boring: the determinism layer",{"_id":718,"publishedAt":719,"slug":720,"sponsored":12,"title":722},"f2b2b7a0-c8e2-4387-8a7d-439b28bab359","2026-10-07T20:03:47.355Z",{"_type":10,"current":721},"implementing-a-modular-master-agent-telemetry-and-diagnostic-framework-in-python-prime-sentinel-command-psc","Implementing a Modular Master-Agent Telemetry & Diagnostic Framework in Python: Prime-Sentinel Command (PSC)",{"_id":724,"publishedAt":725,"slug":726,"sponsored":12,"title":728},"b1af6214-1a59-4309-8484-920746d31d04","2026-10-06T14:00:00.000Z",{"_type":10,"current":727},"the-results-of-the-2026-developer-survey-are-here","The results of the 2026 Developer Survey are here!",{"data":730,"sourceMap":-1},{"count":731,"lastTimestamp":12},0]