Agent Tools Test Report

Branch codex/agent-improvements @ 36803954 · 10 Oct 2026 · For the founder and the branch author

1. Verdict

The retrieval machinery is solid, but the tool the model uses first is weak, and several trust points are unchecked. search and expand enforce filters, composer scope and provenance exactly: 0 filter violations in 1,280 rows checked, 123/124 expand windows byte-identical to direct Turbopuffer reads, 0/3,000 random cases widening scope, and all 632 existing tests pass. The agent finished every one of 105 benchmark and adversarial cases with finishReason=stop. The failures cluster in three places. (1) lookup_catalog resolves only 66% of realistic name and title queries at top-1 and returns nothing for 24%. Apostrophes, honorifics, full names, transliteration variants and famous-vs-obscure namesakes all break it. (2) The agent wiring trusts the model in ways it should not: invented author IDs are accepted, about 2% of citation ids do not resolve and are silently dropped in the UI, and genre filters are never used. (3) No default load path (prod, dev or bench) runs prompt v28, the prompt written for these tools, and v28 itself is out of date. I would not ship this as-is. Most of the high-severity fixes are small or medium in effort.

~5,300checks across 6 test lenses (3,000 are random filter property cases)
53confirmed failures (0 claims failed to reproduce)
8high severity
19medium severity
26low severity
66.2%lookup_catalog top-1 on 601 positive queries
Tool / areaChecksPassedRateHighMedLow
lookup_catalog77654269.8%483
search15713988.5%167
expand27926093.2%006
session / filters3,1593,15299.8%003
end-to-end agent (bench + adversarial)29420670.1%243
prompt alignment13430.8%111
test suite632632100%003

Check counts come from each lens's coverage table. They mix very different units: one property case, one query, one benchmark category. Compare rates within a tool, not across tools. Severities are the verifier's corrected severities.

2. lookup_catalog

Tested through the real tool path (tools.lookup_catalog.execute via createToolsContext, which is the in-memory lookupCatalog plus a live Turbopuffer availability check) against the live catalog: 3,185 authors, 8,359 books, 40 genres. 658 queries had expectations written down before running. On top of those: 101 mechanism probes, 118 independent Turbopuffer count checks and a duplicate-record scan. All read-only.

Coverage

CategorynPassNotes
Author forms (laqab, kunya, nisba, "Imam X", "Ibn X", Arabic with hamza/tashkeel, honorifics)25317067.2% top-1, 72.7% top-5. 72 came back empty: extra tokens, ASCII apostrophes, ties won by obscure namesakes.
Author misspellings127Dawood, Uthaimeen, Katheer, Suyooti, Taimiyah fail. Ghazzali returns the modern Muhammad al-Ghazali.
Ambiguous authors158Ibn Hajar and Ibn Rushd both surface. The classical al-Ghazali, al-Dhahabi, Malik, al-Shatibi and Ibn Hisham lose to namesakes.
Absent authors1414All return [], no false positives.
Edge inputs1910"ibn", "b" and ابن return Ibn al-Imam at 1.0. Two-author and full-lineage queries come back empty.
Hadith collections4026Canonical titles work. Saheeh, Jami at-Tirmidhi, Abu Dawood, an-Nasa'i and "Musnad Imam Ahmad" fail or return derivative works.
Famous books (originals, short names)1348462.7% top-1. Abridgements and supplements outrank originals. Fath al-Bari is missing from the top 5.
Commentary requests107An explicit "Sharh X" usually works. "commentary on X" comes back empty.
Absent books33Correctly empty.
Genres (English / Arabic / transliterated)1238338 empty (tasawwuf, seerah, nahw, fatawa, tajweed...). Words after Arabic و never match.
Absent genres33Correctly empty.
Locale (ar vs en labels)88Arabic labels under ar. Arabic queries resolve under en.
Stretch: English titles and exonyms (not in core rates)241Averroes, Sealed Nectar, Fortress of the Muslim...
indexed flag vs direct Turbopuffer Count118118Metadata consistent for all 62 books checked.

Metrics

Top-1 / top-5 (positive core)398/601 (66.2%) / 428/601 (71.2%)
Returned empty (positive core)146/601 (24.3%); wrong non-empty top-1 57/601
Top-1 by typeauthor 66.8%, book 64.2%, genre 68.0%
Top-1 by scriptArabic 130/163 (79.8%), Latin 268/438 (61.2%)
Queries with ASCII apostrophe/backtick that fail21/23
matchScore calibration1.0: 279 right / 1 wrong. 0.90–0.99: 59/35. 0.80–0.89: 38/19. Only 1.0 is reliable.
Top-1 tied with top-247 cases, 20 wrong (ties broken alphabetically by English label)
Sync lookup latency (warm)p50 1.0 ms, p90 61 ms, p99 265 ms, max 630 ms (blocks the event loop)
Tool execute latency (incl. Turbopuffer)p50 152 ms, p90 503 ms, p99 1.35 s; 0 errors, 0 rate limits at concurrency 3
Output schema validity658/658

Confirmed failures

LC-1All-or-nothing token matching: any extra, honorific, full-name or variant token returns an empty resultHigh

Frequency
146/601 positive core queries returned [] (72 authors, 36 books, 38 genres). Bench: 3/43 live lookups. Adversarial: "Shaykh al-Islam Ibn Taymiyya", فتح الباري ابن حجر and أحكام القرآن ابن العربي all empty.
Expected
Ibn Taymiyya (54), al-Bukhari (215), al-Suyuti (7), Ibn al-Qayyim (14), Ibn Uthaymin (57), al-Fakhr al-Razi (55), al-Risala lil-Shafii at top-1 with a reduced score.
Actual
[] for every input. The model then concludes the source is missing, or drops the filter. In e2e runs it recovers only when Gemini happens to rewrite the name into the catalog's Arabic spelling.
Root cause
packages/api/src/lib/library-index.ts ~264–279: a literal match requires every query token in the candidate. The fuzzy fallback (283–298) requires every token to fuzzy-match within ~20% edit distance. AUTHOR_HONORIFICS covers only imam/shaykh, so رحمه الله, الحافظ, "Hafiz" and "Shaykh al-Islam" stay as tokens. Labels carry only the short laqab. Book lookup never consults author names. No folding of ee/i, oo/u, doubled consonants or iya/iyya.
Fix
Score by IDF-weighted token overlap instead of requiring every token. Strip honorific phrases in both scripts. Add full-name, kunya and laqab aliases for the top ~300 authors. For type=book, also match author names so "Title + Author" works. Normalise Latin transliteration variants.

Verifier: all 13 listed inputs returned [] again, plus "al-Hafiz Ibn Hajar", الحافظ ابن حجر, "Fath al-Bari Sharh Sahih al-Bukhari", genre "Islamic jurisprudence" and "sira". Several earlier misses are now fixed (al-Suyuti, al-Razi, Ibn Uthaymin, "ibn qayyim", "Imam Muslim", creed, Forty Hadith Nawawi). Some failures are pure transliteration variants (Uthaimeen, Saheeh, Riyadh us Saliheen) rather than extra tokens. The 146/601 figure was not re-counted.

Repro
Probe script calling tools.lookup_catalog.execute via createToolsContext with the server env. Inputs: {type:'author', query:'Shaykh al-Islam Ibn Taymiyya'}, ابن تيمية رحمه الله, "Muhammad ibn Ismail al-Bukhari", "Jalal al-Din al-Suyuti", "Ibn Uthaimeen", "Ibn Qayim al-Jawziya", "Fakhr al-Din al-Razi"; {type:'book'} "Saheeh al-Bukhari", "Fath al-Bari Ibn Hajar", الرسالة الشافعي, "Riyadh us Saliheen", "Matn al-Tahawiyya".

LC-2ASCII apostrophe, backtick or curly quote splits a word, so al-Shafi'i, al-Nasa'i, Ibn Sa'd and Zad al-Ma'ad return nothingHigh

Frequency
21/23 queries containing ', `, ‘ or ’
Expected
al-Shafii (20), al-Nasai (49), al-Ashari (58), Ibn Saad (87), al-Sadi (128) and the books.
Actual
[] for all. The U+02BF form ("al-Shafiʿi") works. "Ibn Saʿd" returns al-Layth ibn Sad and other Ibn Sads at 0.933 and misses Ibn Saad (87), because sad ≠ saad.
Root cause
packages/api/src/lib/text-normalize.ts: TRANSLIT_MARKS_RE strips only U+02BB/02BC/02BE/02BF, then TOKEN_SPLIT_RE splits on ASCII '. The possessive rule in library-index.ts only handles a trailing 's.
Fix
Delete ' ` ‘ ’ ʻ–ʿ inside words before splitting (keep the possessive rule). Optionally fold "aa"→"a" in Latin tokens.

Verifier: reproduced for all listed authors and books, and for the backtick, ‘ and ’ forms.

Repro
Probe script, author queries "al-Shafi'i", "al-Nasa'i", "al-Ash'ari", "Ibn Sa'd", "al-Sa'di"; book queries "Zad al-Ma'ad", "Siyar A'lam al-Nubala", "Arba'in Nawawi", "Tafsir al-Sa'di".

LC-3No prominence prior: famous scholars lose to obscure or modern namesakesHigh

Frequency
23 author queries with a wrong top-1 at score ≥0.85. 47 top-1 ties, 20 wrong. Adversarial 7/102 probes. Bench reproduced الغزالي.
Expected
Abu Hamid al-Ghazali 619, Malik ibn Anas 214, Ahmad ibn Hanbal 220, al-Bayhaqi 74, al-Tabari 59, al-Dhahabi 362, Ibn Abd al-Wahhab 152, al-Hakim al-Naysaburi 231 at top-1, or an ambiguity signal.
Actual
al-Ghazali → Muhammad al-Ghazali (d.1416) at 0.9 over 619 at 0.867. "Imam Malik" → Ibn Malik / al-Malik al-Mansur; 214 not in the top 5. "Imam Ahmad" → Ahmad Ahmad Badawi; 220 not in the top 5. al-Hakim → al-Hakim al-Tirmidhi; 231 not in the top 5. al-Tabari, al-Dhahabi and Ibn Abd al-Wahhab lose alphabetical ties to the wrong person.
Root cause
library-index.ts ~271–275: score = 0.8 + 0.2·|query|/|candidate tokens| favours short labels. A stable sort over English labels breaks ties alphabetically. No popularity, indexed-volume or era signal. Honorific stripping turns "Imam Malik" into bare "malik".
Fix
Add a prior (indexed chunk/book count per author) as tiebreak and a small score term. Add aliases "Imam Ahmad"→220, "Imam Malik"/الإمام مالك→214, matched before honorific stripping. Return ties in prior order.

Verifier: every listed case reproduced. The full names "Malik ibn Anas" and "Ahmad ibn Hanbal" resolve at 1.0, so the model only recovers if it retries with the full name. The worst sub-case is the canonical scholar missing from the top 5 (Malik, Ahmad, al-Hakim). Tie cases are softer, because death years let the model disambiguate.

Repro
Probe script, author queries "al-Ghazali", الغزالي, "Imam Malik", "Malik", "Imam Ahmad", "al-Bayhaqi", "al-Tabari", "al-Dhahabi", "Ibn Abd al-Wahhab", "al-Hakim".

LC-4Originals rank below abridgements and supplements; originals with "Sharh" in the title are penalised (Fath al-Bari missing from top 5)High

Frequency
31 book queries with a wrong top-1 at score ≥0.85; at least 16 distinct famous works. Fath al-Bari 4/4 direct probes plus 1/1 e2e.
Expected
Originals first, as docs/agents/chat-retrieval.md promises: Fath al-Bari bi Sharh al-Bukhari, Zad al-Maad (Ibn al-Qayyim), al-Jami li-Ahkam al-Quran, Tahdhib al-Kamal, Muwatta Malik, Ibn Qudama's al-Mughni.
Actual
"Fath al-Bari" / فتح الباري: Ibn Hajar's work scores 0.88−0.12 = 0.76 and is absent from the top 5; the e2e answer cited al-Nukat instead. "Zad al-Maad": the Mukhtasar (0.933) beats the original (0.867). "Tafsir al-Qurtubi" → Durar. "Tahdhib al-Kamal": Ikmal ×2, original 5th. "Bulugh al-Maram" and "Subul al-Salam": derivatives first. الموطأ: none of the three Muwatta records. "al-Mughni" → al-Dhahabi's al-Mughni fi al-Duafa.
Root cause
rankCatalog (library-index.ts ~235–262): commentary markers are only sharh/hashiya/taliq. The 0.12 penalty hits any title containing a marker, including originals and including when the query names that commentary. Mukhtasar, Takmila, Dhayl, Ikmal, Durar, Muntaqa, Nukat and Zawaid are never penalised, and shorter derivative titles also win on the length ratio.
Fix
Penalise only records whose title wraps the queried title (extra words before the query span, or "ala/li/fi" + query). Skip the penalty when the query covers the title's leading tokens. Extend markers in both scripts (مختصر تكملة ذيل إكمال درر). Boost titles that start with the query tokens. Add a الموطأ → Muwatta alias.

Verifier, partly reproduced: everything above reproduced except Siyar. The original is labelled "Siyar Aalam al-Nubala" in English, so "Siyar Alam al-Nubala" fails on aalam≠alam while Takmilat (spelled "Alam") matches at 0.95. That is a label-spelling inconsistency, not derivative ranking. The Arabic سير أعلام النبلاء returns the original at 1.0.

Repro
Probe script, book queries "Fath al-Bari", فتح الباري, "Zad al-Maad", زاد المعاد, "Tafsir al-Qurtubi", "Kitab al-Tawhid", "Bulugh al-Maram", "Tahdhib al-Kamal", "Subul al-Salam", الموطأ, "al-Mughni".

LC-5lookup_catalog fails entirely when Turbopuffer is unavailable, though the catalog match is in memoryMedium

Frequency
1/1 with an invalid Turbopuffer key. A 429 takes the same unguarded path (inferred, not tested).
Expected
In-memory candidates returned with indexed: null (unknown).
Actual
Tool error 401 ... not authorized; the in-memory matches are thrown away.
Root cause
packages/api/src/ai/agentic/tools.ts ~205–212 awaits lookupIndexedSources with no try/catch.
Fix
try/catch, fall back to indexed: null, optionally retry once on 429.

Verifier: medium is fair. If Turbopuffer is fully down, search fails anyway, so the user-facing impact is mostly during throttling.

Repro
Probe script with TURBOPUFFER_API_KEY=invalid, calling lookup_catalog {type:'author', query:'ابن تيمية'}. With the real key it returns 54 (indexed) and 988.

LC-6An exact match on one catalog spelling hides every other spelling variantMedium

Frequency
3 of 4 probed spelling pairs; latent wherever catalog transliteration is inconsistent (yy/y, ae/ai).
Expected
All Tahawiyya commentaries, led by Ibn Abi al-Izz; all Arba'in commentaries.
Actual
"Sharh al-Aqida al-Tahawiyya" → only Khalid al-Muslih at 1.0. "…Tahawiya" → five others at 1.0 (including Ibn Abi al-Izz) without al-Muslih. "Sharh al-Arbain al-Nawawiyya" → one unrelated hit at 0.502, though four "Sharh al-Arbaeen al-Nawawiya" records exist.
Root cause
library-index.ts ~283: fuzzy candidates are considered only when there are no literal hits. No yy→y / ae→ai folding.
Fix
Fold Latin variants in normalizeLookup, and always merge high-quality fuzzy candidates with literal ones.

Verifier: Ibn Abi al-Izz is shown for the single-y spelling and leads for the Arabic شرح العقيدة الطحاوية. He is hidden only for the yy/yyah spellings, which are the most common in English. Related: "taliqat" and "matn" are not markers, so al-Fawzan's Taliqat outranks al-Tahawi's original Matn.

Repro
Probe script, book queries "Sharh al-Aqida al-Tahawiyya", "Sharh al-Aqida al-Tahawiya", "Sharh al-Aqidah al-Tahawiyyah", "Sharh al-Arbain al-Nawawiyya".

LC-7Generic tokens give confident false matches; "Shaykh al-Islam" resolves to al-BazdawiMedium

Frequency
"ibn", ابن, "b", "Abu", "Salih", "a", "i", "Muhammad", "Ahmad"; 3/3 "Shaykh al-Islam" / شيخ الإسلام probes.
Expected
[] or low scores for generic tokens. "Shaykh al-Islam" → Ibn Taymiyya (54) or an explicit ambiguity.
Actual
"ibn"/ابن/"b" → Ibn al-Imam at 1.0; "Abu" → al-Abi at 1.0; "Salih" → Salih Al al-Shaykh at 1.0. "Shaykh al-Islam" → Fakhr al-Islam al-Bazdawi and Zafar al-Islam Khan at 0.867, no Ibn Taymiyya.
Root cause
library-index.ts ~186–203 also normalises catalog labels as author names ("Ibn al-Imam" → "ibn"). canonToken maps b→ibn and abi→abu. "shaykh" and "al" are stripped, leaving "islam", which matches any "…al-Islam" name. Apostrophe splitting (LC-2) leaves one-letter tokens.
Fix
Strip honorifics from queries only; keep the full catalog form as an exact key. Treat multi-word honorifics (shaykh al-islam, hujjat al-islam, شيخ الإسلام) as phrases and alias them. Drop 1-character tokens. Cap scores for queries made only of ibn/abu/al/common first names.

Verifier: all values reproduced. Bare "ibn" or single letters are unlikely model inputs, but "Shaykh al-Islam" is realistic and confidently wrong, which keeps this at medium.

Repro
Probe script, author queries "ibn", "Abu", "Salih", "a", "i", "Shaykh al-Islam", شيخ الإسلام.

LC-8matchScore is miscalibrated; ties hide ambiguityMedium

Frequency
54 of 151 top-1s (36%) in the 0.80–0.99 band are wrong. Only 1.0 is reliable (279/280).
Expected
A high score means a confident identity match; generic queries score low; ambiguity is detectable.
Actual
Book "Sahih", "Tafsir" and كتاب return five arbitrary titles at 0.9. "Bokhari" finds the right al-Bukhari at only 0.6. الرازي gives 0.90/0.90/0.87/0.87/0.87, so the model cannot tell it is ambiguous (feeds AE-4).
Root cause
library-index.ts ~274: the length-ratio formula ignores how informative the query is; ~294 caps fuzzy scores at 0.7·(1−fuseScore).
Fix
Score by matched IDF mass over total query IDF, plus label coverage. Penalise queries of 1–2 common tokens. Expose an ambiguous flag when the top two are within ε.

Verifier: the "54 wrong with score ≥0.85" figure actually counts every wrong top-1 from 0.80 to 0.99. Many of these are ranking failures (LC-3/LC-4) as much as calibration failures.

Repro
Analysis of the 658-query run, plus live queries: book "Siyar Alam al-Nubala", "Sahih", كتاب; author "Bokhari", الرازي, "Tabari".

LC-9Genre coverage is thin, and two English labels are misleadingMedium

Frequency
38/125 genre queries empty. "principles of jurisprudence" and "jurisprudence" map to the wrong genre.
Expected
tasawwuf/zuhd → 23, seerah → 24, nahw → 31, "principles of jurisprudence" → 11 (Usul al-Fiqh). Category 6 labelled e.g. "Hadith Collections".
Actual
Transliterations and English (tasawwuf, Sufism, التصوف, الزهد, seerah, nahw, balagha, fatawa, tajweed, rijal, "legal maxims") return []. "principles of jurisprudence" → 12 "Jurisprudence Principles" (actually legal maxims) at 1.0. "hadith" → category 6, labelled "Sunni Books" (كتب السنة).
Root cause
library-index.ts ~147–164: GENRE_ALIASES has 4 groups. The English labels for categories 6 and 12 are wrong in category_translations.
Fix
One alias row per genre (transliteration + English + Arabic). Relabel 6 "Hadith Collections" and 12 "Legal Maxims and Fiqh Sciences".

Verifier: the Arabic السيرة, النحو and الفتاوى work, so the gap is transliterations and English. There is no dedicated Sufism category: map zuhd/raqa'iq to 23 (Morals and Remembrances) and say so, rather than forcing a mapping.

Repro
Probe script, genre queries "tasawwuf", "Sufism", التصوف, الزهد, "seerah", "nahw", "balagha", "fatawa", "tajweed", "rijal", "legal maxims", "principles of jurisprudence", "hadith"; getCategoryName(6,'en').

LC-10Arabic conjunction و is not split, so genre and title words after it never matchMedium

Frequency
4/7 probed Arabic genre words.
Expected
Genres 20, 30, 12, 26; al-Bidaya wa al-Nihaya; Tarikh al-Tabari.
Actual
[] for genre القضاء, المعاجم, القواعد الفقهية, الطبقات and book البداية النهاية, تاريخ الرسل الملوك.
Root cause
library-index.ts ~127–129 strips only a leading ال; والقضاء stays one token.
Fix
Strip the clitics و ف ب ل before ال on both query and catalog sides.

Verifier: the impact is narrower than "titles containing و". A full title written with و works (البداية والنهاية at 1.0). Only queries that drop the conjunction, or use just the word after it, fail.

Repro
Probe script with the inputs listed under Actual.

LC-11Distinct records with identical English labels are indistinguishableMedium

Frequency
69 same-author identical-label groups (177 books), 64–65 with more than one indexed record.
Expected
Rows the model can tell apart (riwaya or edition), or a canonical one first.
Actual
"Muwatta Malik" → three rows identical in every field at 1.0 (the Abu Mus'ab, Shaybani and Yahya recensions, all indexed). "Talbis Iblis" and "al-Adab al-Mufrad" → two identical rows each. "Nukhbat al-Fikar" → Ibn Hajar's text tied at 1.0 with a modern study of it.
Root cause
English titles drop the Arabic suffix (e.g. - رواية يحيى); the tool output has no variant or edition field.
Fix
Return the Arabic title suffix as variant. Mark a canonical edition, or group duplicates into one candidate with all bookIds.

Verifier: the second "Nukhbat al-Fikar" row is a modern study, not a versification as first reported; a secondary work tying the original supports the finding.

Repro
Duplicate-label scan of the catalog, plus book queries (en) "Muwatta Malik", "Talbis Iblis", "al-Adab al-Mufrad", "Nukhbat al-Fikar".

LC-12Synchronous lookup blocks the event loop up to ~630 ms, plus a strong-consistency Turbopuffer round-trip per callMedium

Frequency
Warm sync p90 61 ms, p99 265 ms, max 630 ms (books p90 180 ms). Execute p50 152 ms, p99 1.35 s.
Expected
Under 10 ms CPU; availability served from memory.
Actual
The Fuse fallback over 8,359 books plus JS Levenshtein runs on the server thread. Misses pay the full fuzzy cost (a made-up title took 297 ms). The availability check alone can cost ~0.6 s (genre "fiqh": 0.1 ms sync, 664 ms total).
Root cause
library-index.ts ~283–298; packages/engine/src/catalog-availability.ts ~55–65 (consistency: strong).
Fix
Precomputed inverted token/trigram index. Cache indexed author/book/genre sets in memory with periodic refresh, or use eventual consistency.
Repro
Probe script timing warm repetitions. Slowest: book "Kitab at-Tawheed Muhammad ibn Abd al-Wahhab" (~630 ms), "Sharh al-Aqida al-Tahawiyya Ibn Abi al-Izz" (~610–629 ms).

LC-13English titles and Latin exonyms are not resolvableLow

Frequency
23/24 stretch cases; 8/8 on re-check.
Expected
Ibn Rushd al-Hafid, al-Raheeq al-Makhtum, Hisn al-Muslim, Riyad al-Salihin, Ihya, Ibn Khaldun, La Tahzan.
Actual
[] for "Averroes", "Avicenna", "The Sealed Nectar", "Fortress of the Muslim", "Gardens of the Righteous", "Revival of the Religious Sciences", "Muqaddimah Ibn Khaldun", "Don't Be Sad". The transliterated forms all resolve.
Root cause
No English-title aliases; BOOK_ALIASES holds only the Nawawi Forty.
Fix
A reviewed alias table for ~100 famous works and figures. Low because the model usually transliterates before calling the tool.
Repro
Probe script with the inputs listed under Actual.

LC-14DB book_versions.is_indexed is false for all 8,583 versions, so lookupCatalog's own indexed field is always falseLow

Frequency
8,583/8,583 versions.
Expected
Flag consistent with Turbopuffer.
Actual
populateDataCache derives record.indexed from the stale flag. The tool masks it by overwriting it with the live Turbopuffer check (118/118 correct).
Root cause
library-index.ts ~450–477 trusts the DB flag. apps/cli/src/index-version.ts does set it on completion, so the production corpus was probably loaded by a path that skipped the update, or the flag was reset later (unverified).
Fix
Backfill the flag from Turbopuffer, or drop the DB-derived field.

Verifier, partly reproduced: the claim that other callers (validator, composer UI) receive the false value is wrong. Only lookupCatalog reads the field, so today it is dead and misleading rather than a live bug. Fixing it alone would not make unindexed filters fail fast (SR-8); that check does not exist.

Repro
Direct read-only DB select counting is_indexed=true (0), compared with the Turbopuffer count checks.

LC-15Some English labels mix Latin and Arabic scriptLow

Frequency
49 book titles and 5 author names (originally reported as 2 seen incidentally).
Expected
Fully transliterated English labels.
Actual
"al-Adhkar li al-Nawawi t Mستو", "Masa'il al-Imam Ahmad wa Ishaq ibn Rahويه", about 10 "Matbu ضمن Muallafat…", broken "-ويه" endings in authors (Ibn Zanjويه, Ibn Miskويه). They reach tool output and the UI.
Root cause
English label generation data.
Fix
Find English labels containing Arabic letters and regenerate them.
Repro
Scan of all English book and author labels for Arabic code points.

The real tools.search.execute with a production createToolsContext, against live Turbopuffer, OpenRouter embeddings and Cohere rerank-v4.0-pro. Inputs were validated with searchInputSchema first, as the AI SDK does. About 180 executions; every returned row (1,280) was checked against the effective filter and the catalog.

Coverage

CategorynPassNotes
Filter semantics (author, book, genre, death-year ranges, AND vs OR, composer intersection, unknown ids, impossible combinations)48480 violations in 560 filtered rows. Impossible combinations rejected before any DB call in 1–12 ms.
Query handling (tashkeel, quotations, English, transliteration, mixed, 1 word, ~2,000 chars, emoji, injection-like)3835All 3 failures are unscoped hadith quotations whose primary collection is missing (SR-4).
Spelling-variant pairs (BM25 top-24 overlap)73Ta marbuta and hamza variants share 0/24 lexical results (SR-6).
Session behaviour (memo, exclusion, history, parallel, failure eviction)118Earlier-turn passages re-delivered; parallel calls lose slots; OR memo key.
Result quality (22 realistic questions, top 20 judged by hand)2219Precision@20 ≈ 0.92. Failures are duplicate-heavy results.
Malformed / edge inputs the model might send1611null for optional args, uppercase UUID, punctuation-only queries.
Latency and payload (sequential)1515All under 2.1 s, exactly 3 network calls each.

Metrics

Rows checked vs filters/catalog1,280 rows, 0 violations
Latency p50 / p90 / max (sequential)1,134 / 1,567 / 2,078 ms. Rerank ≈ 46%, embed ≈ 32%, Turbopuffer ≈ 22%
PayloadTurbopuffer request ≈ 38.5 KB + ≈1.5 KB per earlier call (NotIn list); model-facing output mean 49.5 KB, max 66 KB
Precision@20 (manual, 22 questions)0.92; mean 15.4 distinct books in top 20
Near-duplicate pairs (Jaccard ≥0.5)34 across 22 lists (9 exact repeats within one version)
Primary-source rank, "إنما الأعمال بالنيات…" unscopedSahih al-Bukhari: vector rank 91, BM25 >300. Rank 1 when scoped to the book.
Chunks with null author death year1,847,099 / 7,455,835 (24.8%)
Wasted slots on a related follow-up10/20 re-delivered from the previous turn
Parallel-call slot loss36/40 delivered for two overlapping calls in one step

Confirmed failures

SR-1Search accepts invented author IDs (real IDs of the wrong author) with no provenance checkHigh

Frequency
Bench: 6/63 ID-filtered searches (9.5%) in 4 Arabic cases, 0 in English. Adversarial: 1/78.
Expected
Filter IDs come from lookup_catalog output (the prompt says "Never invent IDs"), or search rejects IDs never resolved in this conversation.
Actual
For al-Ghazali the model used authorIds [505] (al-Mu'ammal ibn Ihab); for al-Qushayri [465] (al-Hazimi). 505 and 465 are those scholars' death years. Also [102] al-Kasani, [262] al-Firyabi ×2, [580] al-Mardawi, [16] Husayn Nassar. All accepted; each returned 10–20 passages from the wrong author, and the intended restriction never applied. In old-classics-ar the Ihya and Qushayriyya were never actually searched, so the answer degraded silently.
Root cause
validateCatalogIds (library-index.ts ~868–891) only checks existence; the session does not track resolved IDs. Model-facing passages omit author_id, so the model has no real ID to copy, and it appears to confuse deathYearAH with an author ID.
Fix
Record lookup_catalog outputs, composer scope and IDs from retrieved rows in RetrievalSession. Reject unseen filter IDs with "resolve X with lookup_catalog", or echo the resolved author names in the search output so the model can notice the mismatch.
Repro
Tool level: search {query:…, filters:{authorIds:[505]}} returns 10 rows, all by المؤمل بن إيهاب; [465] returns 20 rows, all by الحازمي. Only nonexistent IDs ([99999999]) are rejected. E2e: cd apps/cli && bun run bench run --agent <v28-pinned agent> --judge none --case qaf-old-classics-ar (also qaf-calendar-ar, qaf-change-genre-ar, qaf-negative-filter-ar).

SR-2Follow-up turns re-deliver passages already given in earlier turns; the model sees an empty or shortened resultMedium

Frequency
Identical repeat query: 20/20 duplicates. Related query: 10–11/20. Unrelated query: 0/20. Expand: 4/4 neighbours re-delivered on re-expanding a passage expanded last turn.
Expected
Passages still in the 20-message window are excluded from the DB query, so follow-ups spend their 20 slots on new evidence.
Actual
Turbopuffer returns the same passages; projection strips them, so the model sees [] or 9 of 20 and may read that as "no evidence". The UI and DB persist duplicate sources, and the full embed + Turbopuffer + rerank cost is spent.
Root cause
retrieval-session.ts seedHistory fills only sources, never seen; exclusions cover only the current turn. docs/agents/chat-retrieval.md documents this as intended, so it is a design gap, not a code-vs-doc bug.
Fix
In seedHistory, add ids from history tool results within the window to seen (or a separate history-exclusion set). Keep projection dedupe as a safety net. At minimum, return a "no new passages" marker.
Repro
Turn 1: search {query:'فضل الصبر', filters:{authorIds:[54]}}; build history with prepareChatModelMessages; turn 2: same search (20/20 duplicates, model sees 0) or فضل الصبر على البلاء وثوابه (11/20 duplicates).

SR-3Parallel searches in one step can't exclude each other's results, losing up to ~30% of slotsMedium

Frequency
4 parallel near-duplicates: 55 unique vs 80 sequential in one run, 62 vs 80 on re-check (18–25 of 80 slots lost). 2 parallel: 36 vs 40.
Expected
Each call delivers up to 20 new passages. The prompts tell the model to search in parallel.
Actual
Fewer unique passages for the same 4 Turbopuffer and 4 rerank calls.
Root cause
tools.ts ~118–148 snapshots the exclusion filter at call start; seen is updated only on completion, with no backfill.
Fix
Rerank to ~40 and keep the first 20 unseen after a serialised dedupe-and-fill step.
Repro
Queries فضل صلاة الجماعة / … في المسجد / ثواب صلاة الجماعة / أجر صلاة الجماعة via Promise.all on one session, then sequentially on a fresh session.

SR-4Unscoped exact hadith quotations never surface the primary collection, only commentariesMedium

Frequency
Re-check: 0/20 primary rows for 4 of 5 famous matns, 1/20 for the fifth. Same in adversarial e2e quotation cases.
Expected
Sahih al-Bukhari and other containing collections appear in the top 20, so takhrij answers cite the collection.
Actual
Top 20 is commentaries, Arba'in sharhs, lesson transcripts and fatwa sites. E2e answers name al-Tirmidhi or Muslim but cite only secondary passages.
Root cause
packages/engine/src/search.ts: 24 candidates per branch; commentaries quoting the matn dominate both branches; no source-type diversity. Neither the tool description nor the prompt tells the model to scope to the collection for attribution.
Fix
Prompt/tool description: for attribution, resolve the collection and search with bookIds (verified: scoped to Bukhari, the hadith is rank 1). Optionally add a category-6 filtered branch for quote-like queries in the same multiQuery and reserve slots.
Repro
Unfiltered search with إنما الأعمال بالنيات وإنما لكل امرئ ما نوى, من حسن إسلام المرء تركه ما لا يعنيه, لا يؤمن أحدكم…, من صام رمضان ثم أتبعه ستا…; count rows by the eight canonical collectors.

SR-5Duplicate passages from multiple indexed editions and repeated texts crowd result slotsMedium

Frequency
Bench: 275/2,858 delivered rows (9.6%), up to 9/20 in one call. Re-check: 56/160 rows (35%) across 8 queries; book-scoped searches up to 65%, generic unscoped 0–15%. 20/394 books have 2–4 indexed versions. Expand: 5/8 windows had a neighbour ≥40% duplicated from another edition.
Expected
One copy per text per call, one edition per book.
Actual
"علاج العشق" scoped to Ibn al-Qayyim returns the same paragraph from two Zad editions and al-Tibb al-Nabawi (8/20). Bukhari-scoped "إنما الأعمال بالنيات" 13/20 duplicates across four versions. The model may treat two editions as corroborating sources.
Root cause
search.ts and deduplicateEvidence dedupe by chunk id only; all versions are indexed into one namespace.
Fix
Rerank to ~30, collapse rows with shingle Jaccard ≥0.8 or the same content hash, then slice to 20; apply the same to expand neighbours. One-version-per-book alone will not fix it: al-Tibb al-Nabawi is a separate catalog book excerpted from Zad.
Repro
search {query:'علاج العشق', filters:{authorIds:[14]}}; {query:'فضل صيام ست من شوال'} (ranks 1–2 byte-identical); duplicate counting with 3-word shingles, diacritics stripped.

SR-6Hamza/alef and ta marbuta spelling variants give disjoint lexical resultsMedium

Frequency
6/7 pairs have BM25 top-24 overlap of 0–9/24; final top-20 overlap 8–18/20.
Expected
Common user spellings hit the same lexical evidence.
Actual
Disjoint BM25 candidates. Worse, the spellings without hamza or with ه mostly hit manuscript-catalogue and colophon records stored in normalised spelling, not prose on the topic. The vector branch and reranker only partly compensate.
Root cause
search.ts keywordQuery strips only harakat; no أإآ→ا, ة→ه, ى→ي normalisation at index or query time. Alif maqsura on its own does little damage (21/24).
Fix
A normalised FTS attribute at ingestion plus the same function on the query. Interim: run BM25 for both the raw and the normalised query.
Repro
Pairs شروط صحة الصلاة / شروط صحه الصلاه; مسألة خلق القرآن / مسالة خلق القران; أحاديث أبي هريرة… / احاديث ابي هريره…; زيادة الإيمان / الايمان; أركان الإسلام / اركان الاسلام.

SR-7Index/catalog book_id integrity errors: chunks under another book's id, or under an id missing from the catalogMedium

Frequency
Wider than first reported: 2 orphan book_ids (1,381 chunks, both Musnad al-Shafii) and 51 versions (42,221 chunks) stored under a sibling edition's book_id.
Expected
Chunk book_id equals the catalog owner of the version.
Actual
Musnad al-Shafii versions 21495 and 9615 sit under ids not in Postgres: shown as "Unknown Book" and unreachable by a bookIds filter, so ~78% of that book's indexed text cannot be filtered to. Tafsir al-Baghawi version 12217 sits under another id, so lookup reports the catalog book as unindexed. Version 1376 sits under Sahih al-Bukhari's id, which therefore holds 4 versions. Others include Tafsir al-Thaalabi, Ahkam al-Quran lil-Jassas and Bulugh al-Maram. 51 editions wrongly show as unindexed, and citations carry the wrong edition.
Root cause
Ingestion data re-keyed or stale; format-chunk.ts falls back to "Unknown Book".
Fix
Re-key affected versions. Add an integrity job listing Turbopuffer book_ids missing from the catalog or mismatched with the version owner. Resolve the book name via version_id in formatChunk.
Repro
Read-only enumeration of all 8,310 Turbopuffer book_ids (author_id range partitions with group_by), compared with Postgres version owners.

SR-8Filters on catalogued but unindexed books return a silent 0; the author-with-no-books error is misleadingLow

Frequency
4/4 unindexed-book filters (model and composer); 1/1 author with no books.
Expected
An explicit "source not indexed" error, like the impossible-combination rejection.
Actual
0 rows after ~0.5–1 s of embed + Turbopuffer. al-Bazdawi (0 catalog books) is rejected with "No catalog books can match these AND filters… Check book authorship, genre and AH death years", which is true but does not say he has no books.
Root cause
The validator checks the Postgres catalog only and has no indexed check. Both "unindexed" books tested actually have their chunks under a sibling id (SR-7).
Fix
Use cached availability to reject all-unindexed IDs; special-case authors with no books in the message. Low because lookup_catalog already reports indexed:false.
Repro
filters.bookIds=['ab6bbb86-1411-5d7b-830d-8fef2ac5266f'], also 5dd89763… and the same as composer scope; authorIds:[1412].

SR-9deathYearAH filters silently drop the ~25% of the corpus whose authors are living or undatedLow

Frequency
Every deathYearAH call; 24.8% of chunks have no death year.
Expected
"Contemporary" includes living scholars, or the tool says undated/living authors are excluded.
Actual
{gt:1400} returns only authors with a recorded death year; al-Munajjid, Aidh al-Qarni and similar disappear.
Root cause
buildRetrievalFilter adds ['author_death_year','NotEq',null]. Prompt v28 mentions it; the tool description does not, and nothing tells the model to disclose it.
Fix
Document it in the tool description; add an include-undated/living option.
Repro
search {query:'حكم الاحتفال بالمولد النبوي', filters:{deathYearAH:{gt:1400}}} vs unfiltered.

SR-10null for any optional argument fails schema validationLow

Frequency
4/4 null variants rejected. How often Gemini emits null was not measured.
Expected
null treated as omitted.
Actual
"filters: expected object, received null", "filterOperator: Invalid option", etc. Each costs a step.
Root cause
packages/contracts/src/ai/tools.ts schemas use .optional() only.
Fix
z.preprocess to strip null keys, or .nullish() normalised to undefined.
Repro
searchInputSchema.safeParse({query:'الصبر', label:'x', filters:null}); also filterOperator, deathYearAH.gte, bookIds as null.

SR-11Uppercase book UUID rejected as unknownLow

Frequency
1/1 (model filter and composer filter).
Expected
Same result as lowercase (Zad al-Ma'ad, 20 rows).
Actual
"Unknown book ID D179B76B-…; resolve it with lookup_catalog." The schema's z.uuid() accepts it first.
Root cause
validateCatalogIds does an exact-case Map lookup.
Fix
Lowercase ids before validation in both canonicalSearchFilters and the composer path.
Repro
filters.bookIds:['D179B76B-092B-5334-AE83-7DCEFEE1341E'].

SR-12Memo key treats OR-with-one-dimension and AND as different callsLow

Frequency
1/1; needs the model to flip the operator on an otherwise identical query, which is rare.
Expected
Memo hit.
Actual
3 extra network calls and 20 lower-ranked rows (new, because of seen-id exclusion, but not requested).
Root cause
The memo key in tools.ts ~112–118 includes the raw operator.
Fix
Canonicalise the operator to AND when fewer than 2 dimensions are set.
Repro
{query:'التوسل', filters:{authorIds:[14,54]}}, then the same with filterOperator:'OR'.

SR-13Punctuation-only queries pass validation and return 20 content-free dotted-line chunksLow

Frequency
4/4 punctuation-only queries.
Expected
Rejected, or empty without embed/rerank cost.
Actual
20 rows of ". . . . ." chunks at full cost (~1–1.4 s). The index really holds content-free chunks, which can also surface for normal queries containing ellipses.
Root cause
The query schema only does .trim().min(1); ingestion keeps punctuation-only chunks.
Fix
Require /[\p{L}\p{N}]/u (not letters only: digit queries such as hadith numbers are legitimate). Drop low-letter-ratio chunks at ingestion.
Repro
Queries ؟؟؟ !!! ..., *, ..., -.

SR-14Memoised repeats persist full duplicate outputs, and the [search overlap] log measures memo hits, not retrieval overlapLow

Frequency
Scenario with 5 calls and 2 DB calls logged as 60% overlap; otherwise always 0%.
Expected
No repeated stored outputs; overlap metric computed from raw rows.
Actual
The UI stream and persisted message carry repeated copies of the same 20 passages. The model does not see them (projection reduces them to 0), so the cost is payload and storage, not model tokens.
Root cause
retrieval-session.ts run() returns the cached array for memo hits; index.ts computes overlap from deduplicated outputs.
Fix
Return [] or {duplicateOf} for memo hits; compute overlap from raw rows.
Repro
One session: فضل صلاة الجماعة, a whitespace variant, a different label. Calls 2–3 make 0 network calls and return the same 20 ids.

4. expand

Window correctness checked against direct Turbopuffer reads, plus boundaries, exclusion, filter and composer inheritance, the coordinate guard, history seeding, memoisation, chaining, concurrency, multi-edition books, usefulness and live Gemini use.

Coverage

CategorynPassNotes
Middle passages vs direct read3030Exactly s−2..s+2 minus s and delivered ids; byte-identical.
Book start and last chunk1414Books from 42 to 130,456 chunks.
Volume boundaries66Crosses volumes contiguously.
Exclusion of delivered neighbours66
Guessed or modified coordinates refused2020Including version swaps, uppercased UUIDs, other-session coordinates.
Memoisation / chaining / concurrency / schema1616
Filter inheritance from filtered searches2121
Composer-scoped sessions1212
Large exclusion set (300 ids)55110–185 ms.
Coordinates from history129Scope change (EX-1), re-delivery (SR-2), label-less input (EX-4).
Multi-edition books83Expand correct in all 8; 5 had duplicate text from another edition (SR-5).
Usefulness on realistic cases1162 failed upstream (wrong al-Mughni from catalog; search miss).
Page-order integrity (full-book scan)817996,141 chunks; EX-5.
Agent follow-ups with realistic history66Gemini 3.7 Flash, prompt v28.
Agent in-turn expand calls31274 guessed coordinates refused (EX-2).

Metrics

Windows exact vs direct read123/124 (the mismatch is EX-1)
Latencyp50 123 ms, p90 185 ms, max 524 ms
Calls with ≥1 neighbour omitted as already delivered39/124; 6/124 returned [] for that reason alone
Fabricated coordinates refused20/20; 1 false refusal (EX-4)
Books with more than one indexed version20/394 (Bukhari 4, Tafsir Ibn Kathir 4, Abu Dawud 3)
Agent in-turn refusal rate (guessed seq)4/31 (13%)
Sequence gaps or duplicates0 in 394 books

Confirmed failures

No high or medium findings. expand itself is correct; its failures are edge cases and inherited data problems (see also SR-2 and SR-5, which affect expand too).

EX-1Expanding a history passage outside the current composer scope silently returns []Lowwas medium

Frequency
Deterministic, but only when the user changes scope mid-conversation and then asks for context around an old passage.
Expected
An explanatory error, as search gives. Returning out-of-scope neighbours would break the user's scope, so blocking is correct; the silent empty result is the defect.
Actual
[] with no error. It also happens for passages from unfiltered history searches, because seedHistory rebuilds the stored filter with the current composer.
Root cause
seedHistory uses the current composer; the expand path (tools.ts ~165–188) never runs validateRetrievalFilters.
Fix
Check the passage's book against the current scope and throw an explanatory error.
Repro
History search scoped to Sahih al-Bukhari delivers a passage; new turn composer = Sahih Muslim; expand that passage → []. Controls with no composer or composer still on Bukhari return 4 neighbours.

EX-2The model guesses neighbour sequence numbers instead of copying them; the guard rejects correctly but each guess costs a stepLow

Frequency
Expand lens 4/31 in-turn calls (13%); bench 2/11; adversarial 3/12. 0/9 in follow-ups with realistic history.
Expected
Only delivered coordinates are expanded.
Actual
"Expand only coordinates from a retrieved passage…", then a retry. About +1 step and 3–5 s per guess. The model also mixed version ids between editions.
Root cause
The model must copy three coordinates and reaches ±2–3 to read further; the window is fixed at 2 per side.
Fix
Accept the passage id and resolve coordinates from session sources; add a before/after direction, or document chaining in the description.
Repro
Saved agent traces; bench case qaf-book-en (seq 916 requested between delivered 914 and 918).

EX-3Empty or partial expand results give no reason; omitted delivered neighbours look like missing onesLow

Frequency
6/124 returned [] only because of exclusion; 39/124 omitted some; bench 1/11.
Expected
The output notes omitted neighbours ("already delivered: seq 1, 2").
Actual
Bare []; the model made another expand. The description does not mention exclusion. Separately, an identical repeat expand returns the memoised result again as duplicates.
Root cause
Expand uses the session's exclusion filter and dedupes silently.
Fix
Return alreadyDelivered metadata, or document it in the description.
Repro
Bukhari v1284: expand 3268 → [3266, 3267, 3269, 3270]; expand 3270 → [3271, 3272]; expand 3269 → [].

EX-4History passages become unexpandable if the saved search input fails the current schema; the error message is wrongLow

Frequency
1/1, but almost unreachable today: stored inputs were validated by this schema when made, so it needs a future schema change.
Expected
Expansion with composer-only filters, or a truthful error.
Actual
"…Search again if a historical passage has no sequence number", though it has one.
Root cause
seedHistory registers sources only when searchInputSchema.safeParse(part.input) succeeds.
Fix
Register sources with the composer filter, or parse filters leniently.
Repro
History search input {query:'حديث الدين النصيحة'} with no label (or label ''); expand its 4th result.

EX-5Rare chunk-order and page-label anomaliesLow

Frequency
3 boundaries in 96,141 chunks scanned, plus one spliced tail.
Expected
Sequence order equals book order, with correct page labels.
Actual
al-Mughni fi al-Du'afa v5836: after the colophon come entries from the middle of the book (pages 2:851→2:2391), which is genuinely non-contiguous text. Two others are wrong page labels on continuous text (1:72→1:63, 1:208→1:205), which mislead citations. One is uncertain.
Root cause
Ingestion chunk sequence and page assembly.
Fix
An ingestion check for backward or large page jumps and for text after colophons; re-ingest affected books.
Repro
Direct reads of v5836 seqs 591–593, v7662 seqs 54→55 and 173→174, v30920 seqs 703→704.

EX-6Page-level footnotes repeat across neighbouring chunks, inflating expand output and confusing markersLow

Frequency
Corrected: 18/72 random windows (25%); about 100% of windows in page-footnoted editions such as Bukhari v1284 and Tafsir Ibn Kathir v23604 (25–61% of footnote characters duplicated).
Expected
Only the footnotes a chunk references, or dedupe within a window.
Actual
Bukhari seqs 3268 and 3269 carry the same page-392 footnote block, which belongs to seq 3267; 3269 then has a second, different "(١)". Up to ~7k wasted characters per expand, and markers can point to the wrong note.
Root cause
Chunk footnotes hold whole-page footnotes; packages/api/src/ai/format-chunk.ts passes them through, and expandChunk does no dedupe.
Fix
Filter footnotes to markers present in the chunk text, or strip page footnotes already delivered.
Repro
Direct read of Sahih al-Bukhari v1284 seqs 3267–3269; random seed±2 windows across 9 editions.

5. Session and filter machinery

Mocked-model streamAgenticChat scenarios using the real tools, prepareStep, tool context and projection (live Turbopuffer reads), plus property tests of the filter builder and validator against the full catalog.

Coverage

CategorynPassNotes
Mocked-model scenarios3025Failed: step limit gives no answer (AE-1), stale composer id (SF-1), follow-up re-delivery (SR-2), Turbopuffer-down lookup (LC-5), parallel slot loss (SR-3).
Filter builder / validator property tests3,0003,0000 cases widened past composer scope; 0 validator/filter disagreements; memo keys stable under permutation.
Hand-written edge cases2321The 2 failures are catalog issues (LC-4 Zad al-Maad, LC-9 "Sunni Books").
Exclusion list scale (NotIn up to 8,000 ids)66~750 ms flat; excluded ids never return.
Validator vs indexed availability100100All sampled authors are indexed.

Metrics

Runaway turn (mock keeps searching)20 steps, 400 passages, 0 answer characters, finishReason tool-calls
Prompt characters in that turn5,576,866 total; 557,327 at step 20 (messages only)
Parallel vs sequential (4 near-duplicates)55 vs 80 unique passages
Model-correctable errors sent to Sentry6/6 execute-time errors

Confirmed failures

SF-1A stale composer id makes every search in the turn fail, and the model can't see or remove itLowwas medium

Frequency
Every search in an affected turn, but rare: reopening a chat re-hydrates the composer through resolveFilters, which drops unknown ids. Realistic triggers are a tab left open across a catalog change, old clients, or direct API and bench callers.
Expected
Unknown composer ids dropped consistently (and the user told); search continues in the valid scope.
Actual
"Unknown book ID 00000000-…; resolve it with lookup_catalog." on every search, even with the model's own valid filters. The scope instruction shown to the model omits the id, so it is told to resolve an id it never saw.
Root cause
validateCatalogIds(composer) throws (library-index.ts ~797), while resolveFilters (~662) and createToolsContext (tools.ts ~218–223) silently drop unknown ids.
Fix
Normalise composer filters once per request (drop unknown ids, notify the client); throw only for model-supplied ids.
Repro
Composer {bookIds:['00000000-0000-4000-8000-000000000000'], authorIds:[54]}; search {query:'التوحيد'}. Also unknown authorId 99999999 and categoryId 99999.

SF-2Model-correctable tool errors go to Sentry as exceptions; the impossible-filter message claims a user scope that doesn't existLow

Frequency
Every execute-time validation error and every expand guard refusal. Input-schema failures are not captured (they never reach execute). 2/2 impossible-filter messages without a composer.
Expected
Validation failures are not Sentry exceptions; the message reflects whether a scope exists.
Actual
captureException with errorType: tool_call_failed. Message: "…within the user-selected source scope… user-selected scope: {} ([])".
Root cause
index.ts onToolExecutionEnd (~113–127); library-index.ts ~837–843.
Fix
A typed RecoverableToolError that Sentry skips or logs at info; omit the scope clause when the composer is empty.
Repro
No composer; search {filters:{authorIds:[54], bookIds:['f66e9fd8-0e7a-5185-857b-020567db2646']}} with Sentry mocked.

SF-3Docs out of date: benchmarks.md lists a retrieval-call budget, plan.md points to missing notes, bench checks keep a dead error stringLow

Frequency
3 statements.
Expected
Docs match the code (only isStepCount(20)), and references exist.
Actual
docs/agents/benchmarks.md:37 says "retrieval-call budget", contradicting chat-retrieval.md:15. plan.md:3 says release and managed-prompt notes are in chat-retrieval.md; they are not. apps/cli/src/bench/checks.ts:312 special-cases "Retrieval budget exhausted.", which nothing produces any more.
Root cause
Docs not updated with commit 048d2977 (budget removal).
Fix
Remove the budget mention and the dead string; add a prompt-update and label-promotion checklist to chat-retrieval.md.

Verifier: benchmarks.md:39 (qaf-v2 runs the production label) accurately describes the code. The real issue there is a bench-config risk, covered by PR-1.

6. End-to-end agent

Two lenses. (a) The 70-case curated benchmark (92 turns) with the current harness and prompt v28 pinned, plus a 28-case comparison against the production prompt v24. (b) 35 new adversarial cases (17 Arabic, 18 English) with v28, read by hand. Both ran Gemini 3.7 Flash through Google AI Studio, because the production gateway route was unavailable locally (see limitations).

Coverage

LensCategorynPassNotes
BenchCompletion / stability70700 step-cap hits, max 8 steps, 0 empty answers.
BenchCatalog resolution and ambiguity1814All 4 ambiguity cases resolved silently (AE-4).
BenchGenre filtering400 genre lookups, 0 categoryIds filters (AE-3).
BenchChronology128AE-7, AE-8.
BenchFilter-ID provenance63576 invented IDs (SR-1).
BenchExpand calls118EX-2, EX-3.
BenchCitation resolution (turns)9286AE-2.
BenchInsufficient evidence / grounding84AE-9.
BenchPaired comparison vs v242828v28: 24% fewer searches, same cost.
Adv.Unusual author forms66Passed only because Gemini rewrote names into canonical Arabic first.
Adv.Books (commentary / matn / nonexistent)43Fath al-Bari not found (LC-4).
Adv.Ambiguous name21AE-6.
Adv.Two-author comparison22
Adv.Date filters (CE, Hijri century)44
Adv.Exact quote with tashkeel / verse44Primary collections absent (SR-4).
Adv.Expand (long hadith)22Ka'b ibn Malik stitched from 5 chained expands.
Adv.Composer scope conflicts44
Adv.Follow-up needing new evidence22
Adv.Very broad question (step pressure)22Self-scoped within 2–4 steps.
Adv.English question, Arabic-only concept11
Adv.Author not in corpus / not indexed21AE-5.
Adv.Every citation id resolves3523AE-2.
Adv.Direct lookup_catalog probes10281See section 2.
Adv.Direct search/expand/session probes2415See sections 3–4.

Metrics

CompletionBench 70/70 (92 turns), adversarial 35/35; all finishReason=stop; no 429s
Steps per turnBench mean 2.67, max 8; adversarial mean 3.3, max 7
Tool calls (bench)lookup_catalog 43, search 151 (74 unfiltered, 63 ID-filtered, 15 deathYearAH), expand 11; OR used 0 times
Turn latency (bench)p50 33.6 s, p95 57.0 s; first visible text p50 18.1 s
Input tokens per turnmean 59.9k, p95 195.6k; cache-read share 41%
Cost per turn (computed from tokens)with cached input at 10%: mean $0.043, p95 $0.084; 70-case run $3.96 (rerank not priced)
Judge means (0–4, unreviewed judge)grounding 3.83, attribution 3.97, completeness 3.84, uncertainty 3.90
Unresolved citation idsBench 19/1,120 (1.7%) in 6/92 turns; adversarial 13/566 in 12/35 cases
Cross-edition near-duplicate rows275/2,858 (9.6%)
v28 vs v24 (28 cases)searches 57 vs 75; latency p50 34.3 s vs 38.7 s; cost ~$0.055 per turn both

Confirmed failures

AE-2Answers cite passage ids that were never delivered, and nothing validates themHigh

Frequency
Adversarial 12/35 cases, 13 unique bad ids. Bench 6/92 turns, 19/1,120 ids; date-boundary-en lost 12/12. Intermittent: a live re-run of 8 trials had 0 bad ids. Roughly 2% of ids, 1–13% of turns.
Expected
Every citation id is a delivered passage id, or bad ids are repaired or dropped deliberately before reaching the user.
Actual
Of the 13 unique bad ids: 5 are book_ids copied from a delivered row (none were invented from nothing), 3 are truncated UUIDs, the rest have 1–2 segments altered or spliced. One bench turn cited book_id:seq pairs, which crashed the grounding judge. In the web UI (apps/web/src/components/chat/custom-components.tsx ~52–62) unresolved ids are silently dropped, so users see fewer citation pills or none: date-boundary-en showed no citations at all.
Root cause
citation-output-middleware.ts only repairs tag syntax. Rows expose both id and book_id, and prompt v28 says to cite "the result's ids field". Long UUIDs are error-prone to copy.
Fix
Validate ids against session-delivered and history ids; map near-misses (book_id + seq, prefix/suffix) and drop the rest. Better: short ordinal handles mapped server-side. Fix the prompt to id. Add a bench check that fails on unresolved ids.
Repro
Parse citation tags in each saved answer and look each id up in that trial's tool outputs plus history. Bench: cd apps/cli && bun run bench run --agent <v28-pinned agent> --judge none --case qaf-date-boundary-en.

AE-3The agent never uses genre lookup or categoryIds; genre requests run unfiltered with book names in the queryHigh

Frequency
4/4 genre cases; 0 genre lookups and 0 categoryIds filters in 61 lookups and 283 searches across all v28 and v24 runs. A live re-run (2 trials) did the same.
Expected
lookup_catalog {type:'genre', query:'التفسير'} → 3 (score 1.0), then categoryIds:[3] or resolved mufassir authorIds.
Actual
Unfiltered searches such as "تفسير الطبري إياك نعبد…". Only 2/7 searches naming a mufassir returned that tafsir (al-Tabari 0/2, al-Razi 0/2), though both books are indexed. Results were dominated by magazines, lessons and later works; source_scope scored 0.
Root cause
Prompt v28 mentions genres only in passing; nothing tells the model that "in tafsir works" is a genre restriction. Small catalog gap too: كتب التفسير returns [].
Fix
A prompt rule and example (genre phrase → type=genre → categoryIds). Optionally warn when the query text names a catalog author or book and no filter is set.
Repro
cd apps/cli && bun run bench run --agent <v28-pinned agent> --judge none --case qaf-genre-en (also genre-ar, change-genre-ar/en).

AE-1A turn still calling tools at step 20 ends with no answer text ("Interrupted")Mediumwas high

Frequency
1/1 runaway mock scenario (and 1/1 with preambles: only preambles shown). The real model never got near the limit: 0/129 live turns, max 8 steps.
Expected
A grounded answer or closing message before the step limit, e.g. tools disabled on the final step.
Actual
20 steps, finishReason: tool-calls, empty text, 400 passages delivered, prompt grew from 401 to 557,327 characters (5.58M across the turn). classifyChatTurn marks it interrupted; the web shows "Interrupted" with a Retry that replays everything.
Root cause
index.ts sets stopWhen: isStepCount(20); prepareRetrievalStep never sets activeTools or toolChoice: 'none'. Prompt v28 tells the model the system will disable tools after eight calls, which no longer happens (PR-2).
Fix
On the last step return activeTools: [] / toolChoice: 'none'; add an e2e test asserting non-empty text at the limit (TS-1). Cheap.
Repro
Mocked streamAgenticChat where the model emits search {query:'الصلاة N'} every step.

AE-4Ambiguous titles resolved silently (al-Risala → al-Qushayriyya)Medium

Frequency
Silent resolution in 4/4 saved and 4/4 live runs. Strict failures: the ambiguous-book cases (2/4 saved, 2/2 live), whose notes say "Do not silently select". The ambiguous-author cases allow "resolve or disclose", so choosing Fakhr al-Din al-Razi is acceptable.
Expected
Disclose the ambiguity or ask.
Actual
"al-Risala" → al-Qushayriyya, with the EN answer claiming it is "commonly known simply as al-Risala"; the usual referent is al-Shafi'i's.
Root cause
The catalog hides the main candidate: bare الرسالة returns five titles tied at 0.9 without al-Risala lil-Shafi'i, and الرسالة الشافعي returns []. Score ties (LC-8), plus a prompt line ("resolve them using the request's context or ask") read as permission to pick.
Fix
Prompt: disclose when top candidates are within ~0.05 and context doesn't decide. Fix catalog ranking (LC-1, LC-3, LC-8).
Repro
--case qaf-ambiguous-book-ar (also qaf-ambiguous-author-en) with the v28-pinned agent.

AE-5Requested work not in the corpus (Ibn Arabi's al-Futuhat): restriction silently dropped, secondary sources presented as the FutuhatMedium

Frequency
1/2 not-in-corpus cases (single saved trace).
Expected
Say that al-Futuhat is not in the library and label secondary summaries as such (a v28 rule).
Actual
Five lookups found nothing. The answer opens "In his magnum opus, al-Futuhat al-Makkiyya…" with a section of "Key Statements… in al-Futuhat", all cited to encyclopedias, refutations and magazines. Citation pills show the secondary book names, which reduces but does not remove the misleading framing. One citation id was truncated.
Root cause
lookup_catalog returns a bare [] with no not-found signal; the prompt rule is unenforced. Also a trap: the top match for ابن عربي is a different, minor Ibn al-Arabi (d.617) at 0.933.
Fix
A structured no_match / not_indexed result. Prompt: the first sentence must state unavailability; quotes attributed "as quoted by X".
Repro
Adversarial case adv-unindexed-en (asks about Ibn Arabi's al-Futuhat al-Makkiyya).

AE-8"Older classical books" (Arabic) applied no date filter and cited modern authorsMedium

Frequency
No filter plus modern citations in 2/3 Arabic runs; the cutoff was disclosed in 0/3. The English case was correct.
Expected
Choose and disclose a defensible AH cutoff via deathYearAH.
Actual
6 searches without deathYearAH (2 with invented IDs, SR-1); cited authors who died in 1083, 1428 and 1440 AH. One re-run used a silent {lte:800}.
Root cause
The Arabic phrase is treated as a style hint; no tool or prompt nudge.
Fix
An Arabic prompt example; a bench check that cited death years respect the disclosed cutoff.
Repro
--case qaf-old-classics-ar with the v28-pinned agent.

AE-6Ambiguous-name answer adds an uncited section from general knowledgeLowwas medium

Frequency
2/4 runs of the case.
Expected
Each reading cited or marked unavailable.
Actual
For "ماذا قال ابن العربي في تفسير قوله تعالى: ليس كمثله شيء؟", the Abu Bakr Ibn al-Arabi section is cited, but the "if Muhyi al-Din is meant" section describes Fusus al-Hikam positions with no citation. Other runs cited secondary critiques, which is acceptable.
Root cause
Model behaviour against v28 grounding rules. The catalog does signal absence ([] for محيي الدين بن عربي); the model writes from memory anyway.
Fix
A prompt rule for ambiguous alternatives; a judge criterion flagging long uncited paragraphs. Low: the content is broadly accurate and framed as conditional.

AE-7Hijri boundary off by one, and a CE conversion falsely claimed to be conservativeLowwas medium

Frequency
v28: 3 of 8 explicit-boundary cases. "before 300 AH" became lte 300 in 3/3 English runs (Arabic used lt 300, correct). "before 1200 CE" became lte 596 in 5/6 runs. v24 also failed 3 cases.
Expected
"before 300 AH" → {lt:300}; "before 1200 CE" → {lte:595} with disclosure.
Actual
lte 596, while the answer claims this guarantees no author who died after 1200 CE, and the same answer notes that 596 AH runs Oct 1199 to Sep 1200.
Root cause
Model arithmetic and inclusive-boundary errors; the prompt doesn't map before/after to lt/gt.
Fix
Prompt rules for lt/gt/gte/lte and CE→AH conversion; have the tool echo "Applied: death year < N AH". Low exposure: no run retrieved a violating row, and only one catalog author sits on each boundary.
Repro
--case qaf-unknown-dates-en, qaf-calendar-en, qaf-calendar-ar.

AE-9History-window answers deny a message outside the window; absence overclaimed for an invented quoteLow

Frequency
History-window 6/6 runs; invented-quote 2/2 reruns.
Expected
"The start of the conversation is not visible to me"; "not found in the searched corpus".
Actual
"No code word was mentioned in this conversation." For the quote: "There is no source for this quote in any classical Islamic literature". The reasoning there (anachronism) is sound, so it is misleading in method rather than outcome.
Root cause
packages/api/src/ai/agentic/history.ts slices to the last 20 messages and adds no truncation note.
Fix
Add a system note when history is truncated; stronger abstention wording.

7. Prompt alignment

Langfuse platform/main-gemini versions 24, 28 and 29 compared with the tool schemas and code (13 checks, 4 pass). Correct in v28: no search mode, the expand rules, history coordinates and author-filter semantics.

Load pathLabelVersion loadedWritten for these tools?
Production serverproductionv24No (old search {mode} tools, no lookup_catalog)
Dev server (NODE_ENV ≠ production)latestv29No (v24 derivative)
Bench qaf-v1 / qaf-v2productionv24No
Admin playground (manual pick)retrieval-v2v28Yes, but stale (PR-2)

Confirmed failures

PR-1No default path loads the prompt written for these tools: prod and bench load v24, dev loads v29High

Frequency
3/3 default load paths.
Expected
Production and the qaf-v2 bench run the retrieval-v2 prompt (v28 or a corrected successor).
Actual
v28 carries only the retrieval-v2 label and nothing in code pins it. v24 still tells the model to retry with "semantic vs keyword search"; the schema accepts mode:'keyword' and drops it without error. The new tool descriptions partly compensate (they say to use lookup_catalog first). Bench results for this branch measure the new tools under v24 unless the version is pinned.
Root cause
packages/api/src/ai/agentic/index.ts:67 calls getPrompt('platform/main-gemini') with no options; langfuse-client.ts ~76–91 picks production/latest. apps/cli/src/bench/qaf.ts ~181–185 uses production unless promptVersion is set, and agents/qaf-v2.ts sets none.
Fix
Pin promptVersion in agents/qaf-v2.ts; move the production label in the same release as the code (or request label retrieval-v2); log the loaded prompt version.
Repro
Read-only GET of platform/main-gemini by label production (v24), latest (v29) and retrieval-v2 (v28).

PR-2Prompt v28 contradicts the code: "kind", the removed 8-call budget and tool disabling, "omitted to fit context", the "ids" fieldMedium

Frequency
Static; every v28 turn (lines 27, 84, 86, 87, 91, 92).
Expected
Prompt matches tools.ts and contracts: type, no budget or cutoffs, row field id.
Actual
Line 86 "Use lookup_catalog with kind": {kind:'author',…} fails "type: Invalid option" (in the live bench Gemini followed the schema, 0 schema errors). Line 92 describes "at most eight content-retrieval calls", tools being disabled and a stop after two no-progress calls; none exists. Lines 84/91 mention a shared or exhausted budget. Line 91 "some passages may be omitted to fit the model context" is mostly false. Line 27 "cite using the result's ids field" is wrong and contributes to AE-2. Line 87 leaves out filterOperator OR.
Root cause
v28 was written before commit 048d2977 and the kind→type rename; plan.md lists the prompt update as pending.
Fix
Publish a corrected retrieval-v2 version: type, id, filterOperator, the 20-step limit (answer before it), the death-year null caveat, collection scoping for attribution, the genre rule, ambiguity disclosure, before/after boundary rules. Remove budget, no-progress and omission language.

Verifier: "no matchScore/indexed/filterOperator guidance" is overstated. The tool descriptions explain them and reach the model on every call. The real problem is that the prompt contradicts the code. The link to AE-1 is plausible but not confirmed.

PR-3Dev default prompt v29 never names lookup_catalog or filterOperator and still mentions an exhausted budgetLowwas medium

Frequency
1 stale statement, 2 omissions.
Expected
The dev prompt describes the current tools.
Actual
v29 = v24 minus "semantic vs keyword", plus 5 filter lines and "An empty result may mean… an exhausted budget". The omissions are covered by the tool descriptions, and "dimensions are ANDed" is the correct default.
Root cause
Label latest moved to a v24 derivative instead of the retrieval-v2 line.
Fix
Same as PR-1: pin or relabel retrieval-v2.

8. Test suite

PackageTestsPassTypecheck
packages/api485485pass
packages/engine3535pass
packages/contracts7575pass
apps/cli3737pass

No skips and no environment failures. The suite covers memo eviction, expand provenance, projection prefix stability and deduplication well. What it misses is listed below.

Confirmed failures

TS-1No test catches the step-limit no-answer case; the existing step-limit test encodes it as passingLow

Frequency
1 test; no production-wiring e2e tests.
Expected
A streamAgenticChat + mock-model test asserting answer text at the step limit.
Actual
retrieval.test.ts (~388–499) uses fixture tools and plain streamText, asserts 20 steps, 29 DB calls and no errors, and never checks for text. streamAgenticChat is only ever stubbed in tests. Stale composer ids and lookup with Turbopuffer down are untested.
Root cause
Test design.
Fix
Add streamAgenticChat tests with MockLanguageModelV4 for step-limit synthesis, stale composer id and Turbopuffer-down lookup.

TS-2Bench checks score correct behaviour as failuresLow

Frequency
source_scope 0 on empty-intersection en and ar (2/2 trials); expansion_provenance 0 on expansion-en.
Expected
Abstention cases pass when no out-of-scope rows are cited; citing an already-delivered neighbour satisfies requireExpansion.
Actual
Empty-intersection scope (Ibn al-Qayyim AND died before 500 AH) fails any run that retrieves anything, though the agent correctly explained the request is impossible. expansion-en turn 2 cited the next passages, already delivered in turn 1, and scored 0.
Root cause
apps/cli/src/bench/checks.ts scopeCheck and expansionCheck.
Fix
Judge cited rows for abstention scope; accept already-delivered neighbours.

Verifier, partly: the third sub-claim (tool_errors failing on recovered guard errors, 68/70) is a legitimate signal of wasted steps, not a false failure.

TS-3Built-in qaf-v1/qaf-v2 bench agents can't run locally; a reused output path crashes with EEXISTLow

Frequency
1/1.
Expected
Runs against production Gemini as documented.
Actual
"Your AI Gateway has authentication active, but you didn't provide a valid apiKey": the local CLOUDFLARE_WORKERS_AI_API_KEY is a placeholder. Re-using --output gives a raw EEXIST … link error.
Root cause
packages/engine/src/models.ts builds llm from that key.
Fix
Document the gateway key or let bench agents take an explicit modelId; show a friendly "output exists" error. Local developer experience only.
Repro
cd apps/cli && bun run bench run --agent qaf-v2 --judge none --case qaf-author-en --output <path>, then run again with the same path.

9. What works well

10. Prioritised fix list

Effort: S under a day, M a few days, L a week or more, or needs re-indexing.

#FixAddressesEffort
1Publish a corrected retrieval-v2 prompt (type, id, no budget language, genre rule, collection scoping for attribution, ambiguity disclosure, lt/gt boundary rules), pin it in the qaf-v2 bench and move the production label with the release.PR-1, PR-2, PR-3, AE-3, AE-4, AE-7, SR-4S
2Catalog normalisation quick wins: strip in-word apostrophes, split the Arabic و/ف/ب/ل clitics, treat multi-word honorifics as phrases, drop 1-character tokens.LC-2, LC-10, LC-7S
3Validate citation ids server-side: map book_id + seq and prefix near-misses, drop the rest, or switch to short ordinal handles. Add a bench check.AE-2M
4Filter-ID provenance: reject author/book IDs not seen in lookup output, composer scope or retrieved rows; echo resolved names in search output.SR-1, AE-8M
5Catalog scoring: IDF-weighted partial token match, Latin transliteration folding, always merge good fuzzy candidates, an ambiguous flag, and match author names for book queries.LC-1, LC-6, LC-8, AE-4M
6Prominence prior (indexed volume per author) as tiebreak and score term, plus reviewed alias tables for top authors, genres and English titles; relabel categories 6 and 12.LC-3, LC-9, LC-13M
7Derivative ranking: penalise only titles that wrap the query, extend markers (mukhtasar, takmila, dhayl, ikmal, durar…), stop penalising originals that contain "Sharh".LC-4S
8Robustness bundle: disable tools on the final step; try/catch the availability check; normalise composer ids once per request; typed recoverable errors kept out of Sentry. Add streamAgenticChat e2e tests for each.AE-1, LC-5, SF-1, SF-2, TS-1S
9Exclude passages already delivered in the history window, and fill parallel calls after a serialised dedupe; return "already delivered" metadata.SR-2, SR-3, EX-3, SR-14M
10Index hygiene: re-key the 51 misfiled versions and 2 orphan ids, collapse near-duplicate and cross-edition rows before slicing to 20, add a normalised Arabic FTS attribute, drop content-free chunks.SR-7, SR-5, SR-6, SR-13, SR-8L

11. Method and limitations

Six independent test lenses ran against branch codex/agent-improvements @ 36803954: lookup_catalog, search, expand, machinery (session, filters, prompt, test suite), the curated benchmark, and an adversarial e2e set. Each lens used probe scripts calling the real tool code with the server environment against live Turbopuffer, embeddings and rerank. Expectations were written before running where possible. Every candidate failure then went to a separate verification pass that re-ran it with its own probe script and could correct the frequency, severity or root cause. All 53 claims reproduced, 10 of them only partly; this report uses the corrected severities and folds in the corrections. All work was read-only: no repo edits and no writes to the database, Turbopuffer or Langfuse.

Appendix: findings that did not reproduce

None. Every confirmed claim reproduced at least in part. Partial reproductions (LC-4 Siyar, LC-14 impact, SR-14 model visibility, EX-4 reachability, EX-5 page labels, SF-1 frequency, AE-4 case notes, AE-9 invented quote, PR-3, TS-2 tool_errors) are corrected in place above.