Agent tools test report

Branch codex/agent-improvements @ 36803954 · Tested 10 Oct 2026 · lookup_catalog, search, expand, session and filters, end-to-end agent, prompt, test suite

Verdict

Not ready to ship as-is. Filtering, scope and expand are exact: 0 filter violations in 1,280 returned rows, 123 of 124 expand windows byte-identical, 0 of 3,000 random cases widened scope, all 632 existing tests pass, and the agent finished all 105 cases.

The problems are elsewhere. lookup_catalog, the tool the model calls first, puts the right result first for only 66% of realistic queries and returns nothing for 24%. The agent accepts IDs and citations from the model without checking them. And no default path loads the prompt written for these tools. Most of the high-severity fixes are small or medium in effort.

By tool

ToolChecksPass rateFailures H / M / LMain problem
lookup_catalog77669.8%4 / 8 / 3Right result first for 66% of realistic queries; empty for 24%.
search15788.5%1 / 6 / 7Accepts author IDs the model made up, and repeats passages from earlier turns.
expand27993.2%0 / 0 / 6Correct; only edge cases such as silent empty results.
session / filters3,15999.8%0 / 0 / 3Filtering is exact; a stale composer id breaks every search in the turn.
end-to-end agent29470.1%2 / 4 / 3About 2% of citation ids don't resolve, and genre filters are never used.
prompt alignment134 / 131 / 1 / 1Production and bench load v24, dev loads v29; v28, written for these tools, is out of date.
test suite632632 / 6320 / 0 / 3Passes, but no test catches a turn that ends with no answer.

About 5,300 checks in total, of which 3,000 are random filter property cases. 53 confirmed failures: 8 high, 19 medium, 26 low. A check can be a query, a property case or a benchmark category, so compare rates within a tool, not across tools.

High-severity failures

LC-1Any extra token (an honorific, a full name or a spelling variant) gives an empty result

146 of 601 realistic author, book and genre queries (24%) returned nothing; top-1 accuracy 66.2%

Problem
"Shaykh al-Islam Ibn Taymiyya" or "Muhammad ibn Ismail al-Bukhari" should find the scholar with a lower score. Instead they return nothing, because every word in the query has to match.
Fix
Score by weighted word overlap, remove honorific phrases in both scripts, and add full-name aliases and transliteration folding.

LC-2An apostrophe splits the word, so al-Shafi'i, al-Nasa'i, Ibn Sa'd and Zad al-Ma'ad return nothing

21 of 23 queries containing ' ` ‘ or ’

Problem
These are common English spellings. The apostrophe is treated as a word break, so the query comes back empty.
Fix
Remove apostrophes and quote marks inside words before splitting the query into words.

LC-3Famous scholars lose to obscure or modern namesakes

23 author queries with a wrong top result at score ≥0.85; 20 of 47 tied top results wrong

Problem
"al-Ghazali" returns the modern Muhammad al-Ghazali ahead of Abu Hamid. For "Imam Malik", "Imam Ahmad" and "al-Hakim", the scholar meant is not in the top 5 at all.
Fix
Use each author's indexed volume as a tie-break and a small score term. Add aliases such as "Imam Ahmad" and "Imam Malik", matched before honorifics are stripped.

LC-4Abridgements and supplements rank above the original; Fath al-Bari is missing from the top 5

31 book queries with a wrong top result at score ≥0.85, covering at least 16 famous works

Problem
The original should come first. Fath al-Bari is pushed down because "Sharh" in its own title triggers the commentary penalty, and the Mukhtasar of Zad al-Maad beats the original.
Fix
Penalise only titles that wrap the queried title, and add more derivative markers: mukhtasar, takmila, dhayl, ikmal, durar.

SR-1search accepts author IDs the model made up: real IDs, but for the wrong author

6 of 63 ID-filtered searches in the benchmark (9.5%), all in Arabic cases; 1 of 78 in adversarial cases

Problem
For al-Ghazali the model sent authorIds [505], his death year, which is the ID of a different scholar. search returned that scholar's passages, and the answer got worse without anyone noticing.
Fix
Record the IDs that lookup_catalog has returned in this conversation and reject any others. Also show the resolved author names in search output.

AE-2Answers cite passage ids that were never delivered, and nothing checks them

12 of 35 adversarial cases; 6 of 92 benchmark turns (19 of 1,120 ids). Roughly 2% of ids

Problem
Every citation should point to a passage that was delivered. The model sometimes cites book ids or truncated and altered UUIDs instead. The web UI drops these without notice, and one answer showed no citations at all.
Fix
Check ids on the server against delivered passages, repair near-misses and drop the rest, or switch to short numbered handles. Add a bench check for unresolved ids.

AE-3The agent never uses genre lookup or genre filters

4 of 4 genre cases; 0 genre lookups and 0 categoryIds filters across 283 searches

Problem
"In tafsir works" should look up the tafsir genre and filter by it. The agent ran unfiltered searches instead, and only 2 of 7 searches naming a mufassir returned that mufassir's tafsir.
Fix
Add a prompt rule with an example: a genre phrase means a genre lookup, then a categoryIds filter.

PR-1No default path loads the prompt written for these tools

3 of 3 default load paths

Problem
Production and the qaf-v2 bench should run the retrieval-v2 prompt. They load v24, which describes the old search modes and has no lookup_catalog. Dev loads v29.
Fix
Pin the prompt version in the qaf-v2 bench agent, move the production label in the same release as the code, and log which version is loaded.

Fix order

  1. Publish a corrected retrieval-v2 prompt (type, id, no budget language, genre rule, collection scoping, ambiguity disclosure, lt/gt boundary rules) and pin it in bench and production.PR-1, PR-2, PR-3, AE-3, AE-4, AE-7, SR-4S
  2. Quick catalog normalisation fixes: remove apostrophes inside words, split the Arabic و ف ب ل prefixes, treat multi-word honorifics as phrases, drop 1-character tokens.LC-2, LC-10, LC-7S
  3. Check citation ids on the server: repair near-misses, drop the rest, or use short handles; add a bench check.AE-2M
  4. Filter-ID provenance: reject IDs the model has not seen in lookup output, composer scope or retrieved rows; show resolved names in search output.SR-1, AE-8M
  5. Catalog scoring: weighted partial word match, transliteration folding, merge good fuzzy matches, an ambiguous flag, author names in book queries.LC-1, LC-6, LC-8, AE-4M
  6. Prominence prior plus reviewed alias tables for top authors, genres and English titles; relabel categories 6 and 12.LC-3, LC-9, LC-13M
  7. Derivative ranking: penalise only titles that wrap the query; extend markers; stop penalising originals that contain "Sharh".LC-4S
  8. Robustness: disable tools on the final step, catch availability errors, clean composer ids once per request, keep recoverable errors out of Sentry; add e2e tests for each.AE-1, LC-5, SF-1, SF-2, TS-1S
  9. Exclude passages already delivered in the history window, dedupe parallel calls before filling slots, and say when neighbours were already delivered.SR-2, SR-3, EX-3, SR-14M
  10. Index hygiene: re-key the 51 misfiled versions and 2 orphan ids, collapse cross-edition duplicates before slicing to 20, add a normalised Arabic FTS field, drop content-free chunks.SR-7, SR-5, SR-6, SR-13, SR-8L

S = under a day, M = a few days, L = a week or more or needs re-indexing.

Other confirmed failures

45 medium and low findings, by tool

lookup_catalog

  • LC-5Medium Fails entirely when Turbopuffer is unavailable, though matches are in memory1/1 with an invalid key
  • LC-6Medium An exact match on one spelling hides all other spelling variants3 of 4 probed spelling pairs
  • LC-7Medium Generic tokens give confident false matches; "Shaykh al-Islam" resolves to al-Bazdawi3/3 "Shaykh al-Islam" probes
  • LC-8Medium matchScore is miscalibrated, and ties hide ambiguity54 of 151 top results scoring 0.80–0.99 are wrong; only 1.0 is reliable
  • LC-9Medium Thin genre coverage; two English genre labels are misleading38/125 genre queries empty
  • LC-10Medium Arabic و is not split, so words after it never match4/7 probed Arabic genre words
  • LC-11Medium Different records with identical English labels can't be told apart69 identical-label groups (177 books)
  • LC-12Medium Synchronous lookup blocks the event loop up to ~630 msp99 265 ms, max 630 ms; execute p99 1.35 s
  • LC-13Low English titles and Latin exonyms (Averroes, Sealed Nectar) don't resolve23/24 stretch cases
  • LC-14Low DB is_indexed is false for every version; the field is dead and misleading8,583/8,583 versions
  • LC-15Low Some English labels mix Latin and Arabic script49 book titles, 5 author names

search

  • SR-2Medium Follow-up turns re-deliver earlier passages, so the model sees short or empty results10–11/20 duplicates on a related follow-up; 20/20 on a repeat
  • SR-3Medium Parallel searches in one step can't exclude each other's results18–25 of 80 slots lost with 4 parallel calls
  • SR-4Medium Unscoped exact hadith quotations surface only commentaries, never the collection0/20 primary rows for 4 of 5 famous matns
  • SR-5Medium Duplicate text from multiple editions crowds result slots275/2,858 bench rows (9.6%); up to 65% in book-scoped searches
  • SR-6Medium Hamza and ta marbuta spelling variants give disjoint keyword results6/7 pairs overlap 0–9 of 24
  • SR-7Medium Chunks stored under another book's id, or an id missing from the catalog51 versions (42,221 chunks) misfiled; 2 orphan ids (1,381 chunks)
  • SR-8Low Filtering on an unindexed book returns a silent 04/4 unindexed-book filters
  • SR-9Low deathYearAH filters silently drop living and undated authors24.8% of chunks have no death year
  • SR-10Low null for an optional argument fails validation4/4 null variants rejected
  • SR-11Low Uppercase book UUID rejected as unknown1/1
  • SR-12Low Memo key treats single-dimension OR and AND as different calls1/1; rare
  • SR-13Low Punctuation-only queries return 20 content-free chunks4/4 punctuation-only queries
  • SR-14Low Memoised repeats store full duplicate outputs; the overlap log measures the wrong thing1 scenario (5 calls, 2 DB calls)

expand

  • EX-1Lowwas medium Expanding a history passage outside the current scope silently returns []Deterministic, only after a mid-conversation scope change
  • EX-2Low The model guesses neighbour sequence numbers; each refused guess costs a step4/31 in-turn calls (13%)
  • EX-3Low Empty or partial results give no reason6/124 empty only because of exclusion; 39/124 partial
  • EX-4Low History passages become unexpandable if a saved input fails the current schema1/1; almost unreachable today
  • EX-5Low Rare chunk-order and page-label anomalies3 boundaries in 96,141 chunks
  • EX-6Low Page footnotes repeat across neighbouring chunks18/72 random windows (25%)

session / filters

  • SF-1Lowwas medium A stale composer id makes every search in the turn failEvery search in an affected turn; rare
  • SF-2Low Recoverable tool errors go to Sentry as exceptionsEvery execute-time validation error
  • SF-3Low Docs still describe the removed retrieval-call budget3 statements

end-to-end agent

  • AE-1Mediumwas high A turn still calling tools at step 20 ends with no answer ("Interrupted")1/1 mock scenario; 0/129 live turns
  • AE-4Medium Ambiguous titles resolved silently (al-Risala → al-Qushayriyya)Silent in 4/4 saved and 4/4 live runs
  • AE-5Medium A work not in the corpus (al-Futuhat): secondary sources presented as the work1/2 not-in-corpus cases
  • AE-8Medium "Older classical books" in Arabic applied no date filter and cited modern authors2/3 Arabic runs
  • AE-6Lowwas medium Ambiguous-name answer adds an uncited section from general knowledge2/4 runs
  • AE-7Lowwas medium Hijri boundary off by one; CE conversion wrongly called conservative3 of 8 explicit-boundary cases
  • AE-9Low Denies a message outside the history window; overclaims absence of a quoteHistory window 6/6 runs; invented quote 2/2

prompt alignment

  • PR-2Medium Prompt v28 contradicts the code ("kind", the removed 8-call budget, the "ids" field)Every v28 turn
  • PR-3Lowwas medium Dev prompt v29 never names lookup_catalog and still mentions a budget1 stale statement, 2 omissions

test suite

  • TS-1Low No test catches the no-answer step-limit case; the existing test passes it1 test
  • TS-2Low Bench checks score correct behaviour as failures2/2 empty-intersection trials; expansion-en
  • TS-3Low Built-in bench agents can't run locally; a reused output path crashes1/1

Method and limits