Branch codex/agent-improvements @ 36803954 · Tested 10 Oct 2026 · lookup_catalog, search, expand, session and filters, end-to-end agent, prompt, test suite
Verdict
Not ready to ship as-is. Filtering, scope and expand are exact: 0 filter violations in 1,280 returned rows, 123 of 124 expand windows byte-identical, 0 of 3,000 random cases widened scope, all 632 existing tests pass, and the agent finished all 105 cases.
The problems are elsewhere. lookup_catalog, the tool the model calls first, puts the right result first for only 66% of realistic queries and returns nothing for 24%. The agent accepts IDs and citations from the model without checking them. And no default path loads the prompt written for these tools. Most of the high-severity fixes are small or medium in effort.
By tool
Tool
Checks
Pass rate
Failures H / M / L
Main problem
lookup_catalog
776
69.8%
4 / 8 / 3
Right result first for 66% of realistic queries; empty for 24%.
search
157
88.5%
1 / 6 / 7
Accepts author IDs the model made up, and repeats passages from earlier turns.
expand
279
93.2%
0 / 0 / 6
Correct; only edge cases such as silent empty results.
session / filters
3,159
99.8%
0 / 0 / 3
Filtering is exact; a stale composer id breaks every search in the turn.
end-to-end agent
294
70.1%
2 / 4 / 3
About 2% of citation ids don't resolve, and genre filters are never used.
prompt alignment
13
4 / 13
1 / 1 / 1
Production and bench load v24, dev loads v29; v28, written for these tools, is out of date.
test suite
632
632 / 632
0 / 0 / 3
Passes, but no test catches a turn that ends with no answer.
About 5,300 checks in total, of which 3,000 are random filter property cases. 53 confirmed failures: 8 high, 19 medium, 26 low. A check can be a query, a property case or a benchmark category, so compare rates within a tool, not across tools.
High-severity failures
LC-1Any extra token (an honorific, a full name or a spelling variant) gives an empty result
146 of 601 realistic author, book and genre queries (24%) returned nothing; top-1 accuracy 66.2%
Problem
"Shaykh al-Islam Ibn Taymiyya" or "Muhammad ibn Ismail al-Bukhari" should find the scholar with a lower score. Instead they return nothing, because every word in the query has to match.
Fix
Score by weighted word overlap, remove honorific phrases in both scripts, and add full-name aliases and transliteration folding.
LC-2An apostrophe splits the word, so al-Shafi'i, al-Nasa'i, Ibn Sa'd and Zad al-Ma'ad return nothing
21 of 23 queries containing ' ` ‘ or ’
Problem
These are common English spellings. The apostrophe is treated as a word break, so the query comes back empty.
Fix
Remove apostrophes and quote marks inside words before splitting the query into words.
LC-3Famous scholars lose to obscure or modern namesakes
23 author queries with a wrong top result at score ≥0.85; 20 of 47 tied top results wrong
Problem
"al-Ghazali" returns the modern Muhammad al-Ghazali ahead of Abu Hamid. For "Imam Malik", "Imam Ahmad" and "al-Hakim", the scholar meant is not in the top 5 at all.
Fix
Use each author's indexed volume as a tie-break and a small score term. Add aliases such as "Imam Ahmad" and "Imam Malik", matched before honorifics are stripped.
LC-4Abridgements and supplements rank above the original; Fath al-Bari is missing from the top 5
31 book queries with a wrong top result at score ≥0.85, covering at least 16 famous works
Problem
The original should come first. Fath al-Bari is pushed down because "Sharh" in its own title triggers the commentary penalty, and the Mukhtasar of Zad al-Maad beats the original.
Fix
Penalise only titles that wrap the queried title, and add more derivative markers: mukhtasar, takmila, dhayl, ikmal, durar.
SR-1search accepts author IDs the model made up: real IDs, but for the wrong author
6 of 63 ID-filtered searches in the benchmark (9.5%), all in Arabic cases; 1 of 78 in adversarial cases
Problem
For al-Ghazali the model sent authorIds [505], his death year, which is the ID of a different scholar. search returned that scholar's passages, and the answer got worse without anyone noticing.
Fix
Record the IDs that lookup_catalog has returned in this conversation and reject any others. Also show the resolved author names in search output.
AE-2Answers cite passage ids that were never delivered, and nothing checks them
12 of 35 adversarial cases; 6 of 92 benchmark turns (19 of 1,120 ids). Roughly 2% of ids
Problem
Every citation should point to a passage that was delivered. The model sometimes cites book ids or truncated and altered UUIDs instead. The web UI drops these without notice, and one answer showed no citations at all.
Fix
Check ids on the server against delivered passages, repair near-misses and drop the rest, or switch to short numbered handles. Add a bench check for unresolved ids.
AE-3The agent never uses genre lookup or genre filters
4 of 4 genre cases; 0 genre lookups and 0 categoryIds filters across 283 searches
Problem
"In tafsir works" should look up the tafsir genre and filter by it. The agent ran unfiltered searches instead, and only 2 of 7 searches naming a mufassir returned that mufassir's tafsir.
Fix
Add a prompt rule with an example: a genre phrase means a genre lookup, then a categoryIds filter.
PR-1No default path loads the prompt written for these tools
3 of 3 default load paths
Problem
Production and the qaf-v2 bench should run the retrieval-v2 prompt. They load v24, which describes the old search modes and has no lookup_catalog. Dev loads v29.
Fix
Pin the prompt version in the qaf-v2 bench agent, move the production label in the same release as the code, and log which version is loaded.
Fix order
Publish a corrected retrieval-v2 prompt (type, id, no budget language, genre rule, collection scoping, ambiguity disclosure, lt/gt boundary rules) and pin it in bench and production.PR-1, PR-2, PR-3, AE-3, AE-4, AE-7, SR-4S
Quick catalog normalisation fixes: remove apostrophes inside words, split the Arabic و ف ب ل prefixes, treat multi-word honorifics as phrases, drop 1-character tokens.LC-2, LC-10, LC-7S
Check citation ids on the server: repair near-misses, drop the rest, or use short handles; add a bench check.AE-2M
Filter-ID provenance: reject IDs the model has not seen in lookup output, composer scope or retrieved rows; show resolved names in search output.SR-1, AE-8M
Catalog scoring: weighted partial word match, transliteration folding, merge good fuzzy matches, an ambiguous flag, author names in book queries.LC-1, LC-6, LC-8, AE-4M
Prominence prior plus reviewed alias tables for top authors, genres and English titles; relabel categories 6 and 12.LC-3, LC-9, LC-13M
Derivative ranking: penalise only titles that wrap the query; extend markers; stop penalising originals that contain "Sharh".LC-4S
Robustness: disable tools on the final step, catch availability errors, clean composer ids once per request, keep recoverable errors out of Sentry; add e2e tests for each.AE-1, LC-5, SF-1, SF-2, TS-1S
Exclude passages already delivered in the history window, dedupe parallel calls before filling slots, and say when neighbours were already delivered.SR-2, SR-3, EX-3, SR-14M
Index hygiene: re-key the 51 misfiled versions and 2 orphan ids, collapse cross-edition duplicates before slicing to 20, add a normalised Arabic FTS field, drop content-free chunks.SR-7, SR-5, SR-6, SR-13, SR-8L
S = under a day, M = a few days, L = a week or more or needs re-indexing.
Other confirmed failures
45 medium and low findings, by tool
lookup_catalog
LC-5Medium Fails entirely when Turbopuffer is unavailable, though matches are in memory1/1 with an invalid key
LC-6Medium An exact match on one spelling hides all other spelling variants3 of 4 probed spelling pairs
LC-7Medium Generic tokens give confident false matches; "Shaykh al-Islam" resolves to al-Bazdawi3/3 "Shaykh al-Islam" probes
LC-8Medium matchScore is miscalibrated, and ties hide ambiguity54 of 151 top results scoring 0.80–0.99 are wrong; only 1.0 is reliable
LC-9Medium Thin genre coverage; two English genre labels are misleading38/125 genre queries empty
LC-10Medium Arabic و is not split, so words after it never match4/7 probed Arabic genre words
LC-11Medium Different records with identical English labels can't be told apart69 identical-label groups (177 books)
LC-12Medium Synchronous lookup blocks the event loop up to ~630 msp99 265 ms, max 630 ms; execute p99 1.35 s
LC-13Low English titles and Latin exonyms (Averroes, Sealed Nectar) don't resolve23/24 stretch cases
LC-14Low DB is_indexed is false for every version; the field is dead and misleading8,583/8,583 versions
LC-15Low Some English labels mix Latin and Arabic script49 book titles, 5 author names
search
SR-2Medium Follow-up turns re-deliver earlier passages, so the model sees short or empty results10–11/20 duplicates on a related follow-up; 20/20 on a repeat
SR-3Medium Parallel searches in one step can't exclude each other's results18–25 of 80 slots lost with 4 parallel calls
SR-4Medium Unscoped exact hadith quotations surface only commentaries, never the collection0/20 primary rows for 4 of 5 famous matns
SR-5Medium Duplicate text from multiple editions crowds result slots275/2,858 bench rows (9.6%); up to 65% in book-scoped searches
SR-6Medium Hamza and ta marbuta spelling variants give disjoint keyword results6/7 pairs overlap 0–9 of 24
SR-7Medium Chunks stored under another book's id, or an id missing from the catalog51 versions (42,221 chunks) misfiled; 2 orphan ids (1,381 chunks)
SR-8Low Filtering on an unindexed book returns a silent 04/4 unindexed-book filters
SR-9Low deathYearAH filters silently drop living and undated authors24.8% of chunks have no death year
SR-10Lownull for an optional argument fails validation4/4 null variants rejected
SR-11Low Uppercase book UUID rejected as unknown1/1
SR-12Low Memo key treats single-dimension OR and AND as different calls1/1; rare
TS-3Low Built-in bench agents can't run locally; a reused output path crashes1/1
Method and limits
Six test lenses called the real tool code against live Turbopuffer, embeddings and rerank. All of it was read-only: no repo, database or prompt writes.
A separate verification pass re-ran every candidate failure. All 53 reproduced, 10 of them only in part. This page uses the verifier's corrected severities.
Agent runs used Gemini 3.7 Flash through Google AI Studio, because the production gateway returned 401 locally. Latency, caching and cost may differ in production.
Each benchmark and adversarial case ran once, so model-behaviour frequencies come from small samples. Adversarial verdicts were read by hand, and nobody reviewed the bench judge's scores.
Not measured: which query forms the live model actually sends to lookup_catalog, how often it sends null arguments, and how the availability check behaves on a 429.