Agent Tools Test Report
1. Verdict
The retrieval machinery is solid, but the tool the model uses first is weak, and several trust points are unchecked. search and expand enforce filters, composer scope and provenance exactly: 0 filter violations in 1,280 rows checked, 123/124 expand windows byte-identical to direct Turbopuffer reads, 0/3,000 random cases widening scope, and all 632 existing tests pass. The agent finished every one of 105 benchmark and adversarial cases with finishReason=stop. The failures cluster in three places. (1) lookup_catalog resolves only 66% of realistic name and title queries at top-1 and returns nothing for 24%. Apostrophes, honorifics, full names, transliteration variants and famous-vs-obscure namesakes all break it. (2) The agent wiring trusts the model in ways it should not: invented author IDs are accepted, about 2% of citation ids do not resolve and are silently dropped in the UI, and genre filters are never used. (3) No default load path (prod, dev or bench) runs prompt v28, the prompt written for these tools, and v28 itself is out of date. I would not ship this as-is. Most of the high-severity fixes are small or medium in effort.
| Tool / area | Checks | Passed | Rate | High | Med | Low |
|---|---|---|---|---|---|---|
| lookup_catalog | 776 | 542 | 69.8% | 4 | 8 | 3 |
| search | 157 | 139 | 88.5% | 1 | 6 | 7 |
| expand | 279 | 260 | 93.2% | 0 | 0 | 6 |
| session / filters | 3,159 | 3,152 | 99.8% | 0 | 0 | 3 |
| end-to-end agent (bench + adversarial) | 294 | 206 | 70.1% | 2 | 4 | 3 |
| prompt alignment | 13 | 4 | 30.8% | 1 | 1 | 1 |
| test suite | 632 | 632 | 100% | 0 | 0 | 3 |
Check counts come from each lens's coverage table. They mix very different units: one property case, one query, one benchmark category. Compare rates within a tool, not across tools. Severities are the verifier's corrected severities.
2. lookup_catalog
Tested through the real tool path (tools.lookup_catalog.execute via createToolsContext, which is the in-memory lookupCatalog plus a live Turbopuffer availability check) against the live catalog: 3,185 authors, 8,359 books, 40 genres. 658 queries had expectations written down before running. On top of those: 101 mechanism probes, 118 independent Turbopuffer count checks and a duplicate-record scan. All read-only.
Coverage
| Category | n | Pass | Notes |
|---|---|---|---|
| Author forms (laqab, kunya, nisba, "Imam X", "Ibn X", Arabic with hamza/tashkeel, honorifics) | 253 | 170 | 67.2% top-1, 72.7% top-5. 72 came back empty: extra tokens, ASCII apostrophes, ties won by obscure namesakes. |
| Author misspellings | 12 | 7 | Dawood, Uthaimeen, Katheer, Suyooti, Taimiyah fail. Ghazzali returns the modern Muhammad al-Ghazali. |
| Ambiguous authors | 15 | 8 | Ibn Hajar and Ibn Rushd both surface. The classical al-Ghazali, al-Dhahabi, Malik, al-Shatibi and Ibn Hisham lose to namesakes. |
| Absent authors | 14 | 14 | All return [], no false positives. |
| Edge inputs | 19 | 10 | "ibn", "b" and ابن return Ibn al-Imam at 1.0. Two-author and full-lineage queries come back empty. |
| Hadith collections | 40 | 26 | Canonical titles work. Saheeh, Jami at-Tirmidhi, Abu Dawood, an-Nasa'i and "Musnad Imam Ahmad" fail or return derivative works. |
| Famous books (originals, short names) | 134 | 84 | 62.7% top-1. Abridgements and supplements outrank originals. Fath al-Bari is missing from the top 5. |
| Commentary requests | 10 | 7 | An explicit "Sharh X" usually works. "commentary on X" comes back empty. |
| Absent books | 3 | 3 | Correctly empty. |
| Genres (English / Arabic / transliterated) | 123 | 83 | 38 empty (tasawwuf, seerah, nahw, fatawa, tajweed...). Words after Arabic و never match. |
| Absent genres | 3 | 3 | Correctly empty. |
| Locale (ar vs en labels) | 8 | 8 | Arabic labels under ar. Arabic queries resolve under en. |
| Stretch: English titles and exonyms (not in core rates) | 24 | 1 | Averroes, Sealed Nectar, Fortress of the Muslim... |
indexed flag vs direct Turbopuffer Count | 118 | 118 | Metadata consistent for all 62 books checked. |
Metrics
| Top-1 / top-5 (positive core) | 398/601 (66.2%) / 428/601 (71.2%) |
|---|---|
| Returned empty (positive core) | 146/601 (24.3%); wrong non-empty top-1 57/601 |
| Top-1 by type | author 66.8%, book 64.2%, genre 68.0% |
| Top-1 by script | Arabic 130/163 (79.8%), Latin 268/438 (61.2%) |
| Queries with ASCII apostrophe/backtick that fail | 21/23 |
| matchScore calibration | 1.0: 279 right / 1 wrong. 0.90–0.99: 59/35. 0.80–0.89: 38/19. Only 1.0 is reliable. |
| Top-1 tied with top-2 | 47 cases, 20 wrong (ties broken alphabetically by English label) |
| Sync lookup latency (warm) | p50 1.0 ms, p90 61 ms, p99 265 ms, max 630 ms (blocks the event loop) |
| Tool execute latency (incl. Turbopuffer) | p50 152 ms, p90 503 ms, p99 1.35 s; 0 errors, 0 rate limits at concurrency 3 |
| Output schema validity | 658/658 |
Confirmed failures
LC-1All-or-nothing token matching: any extra, honorific, full-name or variant token returns an empty resultHigh
- Frequency
- 146/601 positive core queries returned
[](72 authors, 36 books, 38 genres). Bench: 3/43 live lookups. Adversarial: "Shaykh al-Islam Ibn Taymiyya", فتح الباري ابن حجر and أحكام القرآن ابن العربي all empty. - Expected
- Ibn Taymiyya (54), al-Bukhari (215), al-Suyuti (7), Ibn al-Qayyim (14), Ibn Uthaymin (57), al-Fakhr al-Razi (55), al-Risala lil-Shafii at top-1 with a reduced score.
- Actual
[]for every input. The model then concludes the source is missing, or drops the filter. In e2e runs it recovers only when Gemini happens to rewrite the name into the catalog's Arabic spelling.- Root cause
packages/api/src/lib/library-index.ts~264–279: a literal match requires every query token in the candidate. The fuzzy fallback (283–298) requires every token to fuzzy-match within ~20% edit distance.AUTHOR_HONORIFICScovers only imam/shaykh, so رحمه الله, الحافظ, "Hafiz" and "Shaykh al-Islam" stay as tokens. Labels carry only the short laqab. Book lookup never consults author names. No folding of ee/i, oo/u, doubled consonants or iya/iyya.- Fix
- Score by IDF-weighted token overlap instead of requiring every token. Strip honorific phrases in both scripts. Add full-name, kunya and laqab aliases for the top ~300 authors. For
type=book, also match author names so "Title + Author" works. Normalise Latin transliteration variants.
Verifier: all 13 listed inputs returned [] again, plus "al-Hafiz Ibn Hajar", الحافظ ابن حجر, "Fath al-Bari Sharh Sahih al-Bukhari", genre "Islamic jurisprudence" and "sira". Several earlier misses are now fixed (al-Suyuti, al-Razi, Ibn Uthaymin, "ibn qayyim", "Imam Muslim", creed, Forty Hadith Nawawi). Some failures are pure transliteration variants (Uthaimeen, Saheeh, Riyadh us Saliheen) rather than extra tokens. The 146/601 figure was not re-counted.
Repro
tools.lookup_catalog.execute via createToolsContext with the server env. Inputs: {type:'author', query:'Shaykh al-Islam Ibn Taymiyya'}, ابن تيمية رحمه الله, "Muhammad ibn Ismail al-Bukhari", "Jalal al-Din al-Suyuti", "Ibn Uthaimeen", "Ibn Qayim al-Jawziya", "Fakhr al-Din al-Razi"; {type:'book'} "Saheeh al-Bukhari", "Fath al-Bari Ibn Hajar", الرسالة الشافعي, "Riyadh us Saliheen", "Matn al-Tahawiyya".LC-2ASCII apostrophe, backtick or curly quote splits a word, so al-Shafi'i, al-Nasa'i, Ibn Sa'd and Zad al-Ma'ad return nothingHigh
- Frequency
- 21/23 queries containing
',`, ‘ or ’ - Expected
- al-Shafii (20), al-Nasai (49), al-Ashari (58), Ibn Saad (87), al-Sadi (128) and the books.
- Actual
[]for all. The U+02BF form ("al-Shafiʿi") works. "Ibn Saʿd" returns al-Layth ibn Sad and other Ibn Sads at 0.933 and misses Ibn Saad (87), becausesad ≠ saad.- Root cause
packages/api/src/lib/text-normalize.ts:TRANSLIT_MARKS_REstrips only U+02BB/02BC/02BE/02BF, thenTOKEN_SPLIT_REsplits on ASCII'. The possessive rule in library-index.ts only handles a trailing 's.- Fix
- Delete
' ` ‘ ’ ʻ–ʿinside words before splitting (keep the possessive rule). Optionally fold "aa"→"a" in Latin tokens.
Verifier: reproduced for all listed authors and books, and for the backtick, ‘ and ’ forms.
Repro
LC-3No prominence prior: famous scholars lose to obscure or modern namesakesHigh
- Frequency
- 23 author queries with a wrong top-1 at score ≥0.85. 47 top-1 ties, 20 wrong. Adversarial 7/102 probes. Bench reproduced الغزالي.
- Expected
- Abu Hamid al-Ghazali 619, Malik ibn Anas 214, Ahmad ibn Hanbal 220, al-Bayhaqi 74, al-Tabari 59, al-Dhahabi 362, Ibn Abd al-Wahhab 152, al-Hakim al-Naysaburi 231 at top-1, or an ambiguity signal.
- Actual
- al-Ghazali → Muhammad al-Ghazali (d.1416) at 0.9 over 619 at 0.867. "Imam Malik" → Ibn Malik / al-Malik al-Mansur; 214 not in the top 5. "Imam Ahmad" → Ahmad Ahmad Badawi; 220 not in the top 5. al-Hakim → al-Hakim al-Tirmidhi; 231 not in the top 5. al-Tabari, al-Dhahabi and Ibn Abd al-Wahhab lose alphabetical ties to the wrong person.
- Root cause
- library-index.ts ~271–275:
score = 0.8 + 0.2·|query|/|candidate tokens|favours short labels. A stable sort over English labels breaks ties alphabetically. No popularity, indexed-volume or era signal. Honorific stripping turns "Imam Malik" into bare "malik". - Fix
- Add a prior (indexed chunk/book count per author) as tiebreak and a small score term. Add aliases "Imam Ahmad"→220, "Imam Malik"/الإمام مالك→214, matched before honorific stripping. Return ties in prior order.
Verifier: every listed case reproduced. The full names "Malik ibn Anas" and "Ahmad ibn Hanbal" resolve at 1.0, so the model only recovers if it retries with the full name. The worst sub-case is the canonical scholar missing from the top 5 (Malik, Ahmad, al-Hakim). Tie cases are softer, because death years let the model disambiguate.
Repro
LC-4Originals rank below abridgements and supplements; originals with "Sharh" in the title are penalised (Fath al-Bari missing from top 5)High
- Frequency
- 31 book queries with a wrong top-1 at score ≥0.85; at least 16 distinct famous works. Fath al-Bari 4/4 direct probes plus 1/1 e2e.
- Expected
- Originals first, as
docs/agents/chat-retrieval.mdpromises: Fath al-Bari bi Sharh al-Bukhari, Zad al-Maad (Ibn al-Qayyim), al-Jami li-Ahkam al-Quran, Tahdhib al-Kamal, Muwatta Malik, Ibn Qudama's al-Mughni. - Actual
- "Fath al-Bari" / فتح الباري: Ibn Hajar's work scores 0.88−0.12 = 0.76 and is absent from the top 5; the e2e answer cited al-Nukat instead. "Zad al-Maad": the Mukhtasar (0.933) beats the original (0.867). "Tafsir al-Qurtubi" → Durar. "Tahdhib al-Kamal": Ikmal ×2, original 5th. "Bulugh al-Maram" and "Subul al-Salam": derivatives first. الموطأ: none of the three Muwatta records. "al-Mughni" → al-Dhahabi's al-Mughni fi al-Duafa.
- Root cause
rankCatalog(library-index.ts ~235–262): commentary markers are only sharh/hashiya/taliq. The 0.12 penalty hits any title containing a marker, including originals and including when the query names that commentary. Mukhtasar, Takmila, Dhayl, Ikmal, Durar, Muntaqa, Nukat and Zawaid are never penalised, and shorter derivative titles also win on the length ratio.- Fix
- Penalise only records whose title wraps the queried title (extra words before the query span, or "ala/li/fi" + query). Skip the penalty when the query covers the title's leading tokens. Extend markers in both scripts (مختصر تكملة ذيل إكمال درر). Boost titles that start with the query tokens. Add a الموطأ → Muwatta alias.
Verifier, partly reproduced: everything above reproduced except Siyar. The original is labelled "Siyar Aalam al-Nubala" in English, so "Siyar Alam al-Nubala" fails on aalam≠alam while Takmilat (spelled "Alam") matches at 0.95. That is a label-spelling inconsistency, not derivative ranking. The Arabic سير أعلام النبلاء returns the original at 1.0.
Repro
LC-5lookup_catalog fails entirely when Turbopuffer is unavailable, though the catalog match is in memoryMedium
- Frequency
- 1/1 with an invalid Turbopuffer key. A 429 takes the same unguarded path (inferred, not tested).
- Expected
- In-memory candidates returned with
indexed: null(unknown). - Actual
- Tool error
401 ... not authorized; the in-memory matches are thrown away. - Root cause
packages/api/src/ai/agentic/tools.ts~205–212 awaitslookupIndexedSourceswith no try/catch.- Fix
- try/catch, fall back to
indexed: null, optionally retry once on 429.
Verifier: medium is fair. If Turbopuffer is fully down, search fails anyway, so the user-facing impact is mostly during throttling.
Repro
TURBOPUFFER_API_KEY=invalid, calling lookup_catalog {type:'author', query:'ابن تيمية'}. With the real key it returns 54 (indexed) and 988.LC-6An exact match on one catalog spelling hides every other spelling variantMedium
- Frequency
- 3 of 4 probed spelling pairs; latent wherever catalog transliteration is inconsistent (yy/y, ae/ai).
- Expected
- All Tahawiyya commentaries, led by Ibn Abi al-Izz; all Arba'in commentaries.
- Actual
- "Sharh al-Aqida al-Tahawiyya" → only Khalid al-Muslih at 1.0. "…Tahawiya" → five others at 1.0 (including Ibn Abi al-Izz) without al-Muslih. "Sharh al-Arbain al-Nawawiyya" → one unrelated hit at 0.502, though four "Sharh al-Arbaeen al-Nawawiya" records exist.
- Root cause
- library-index.ts ~283: fuzzy candidates are considered only when there are no literal hits. No yy→y / ae→ai folding.
- Fix
- Fold Latin variants in
normalizeLookup, and always merge high-quality fuzzy candidates with literal ones.
Verifier: Ibn Abi al-Izz is shown for the single-y spelling and leads for the Arabic شرح العقيدة الطحاوية. He is hidden only for the yy/yyah spellings, which are the most common in English. Related: "taliqat" and "matn" are not markers, so al-Fawzan's Taliqat outranks al-Tahawi's original Matn.
Repro
LC-7Generic tokens give confident false matches; "Shaykh al-Islam" resolves to al-BazdawiMedium
- Frequency
- "ibn", ابن, "b", "Abu", "Salih", "a", "i", "Muhammad", "Ahmad"; 3/3 "Shaykh al-Islam" / شيخ الإسلام probes.
- Expected
[]or low scores for generic tokens. "Shaykh al-Islam" → Ibn Taymiyya (54) or an explicit ambiguity.- Actual
- "ibn"/ابن/"b" → Ibn al-Imam at 1.0; "Abu" → al-Abi at 1.0; "Salih" → Salih Al al-Shaykh at 1.0. "Shaykh al-Islam" → Fakhr al-Islam al-Bazdawi and Zafar al-Islam Khan at 0.867, no Ibn Taymiyya.
- Root cause
- library-index.ts ~186–203 also normalises catalog labels as author names ("Ibn al-Imam" → "ibn").
canonTokenmaps b→ibn and abi→abu. "shaykh" and "al" are stripped, leaving "islam", which matches any "…al-Islam" name. Apostrophe splitting (LC-2) leaves one-letter tokens. - Fix
- Strip honorifics from queries only; keep the full catalog form as an exact key. Treat multi-word honorifics (shaykh al-islam, hujjat al-islam, شيخ الإسلام) as phrases and alias them. Drop 1-character tokens. Cap scores for queries made only of ibn/abu/al/common first names.
Verifier: all values reproduced. Bare "ibn" or single letters are unlikely model inputs, but "Shaykh al-Islam" is realistic and confidently wrong, which keeps this at medium.
Repro
LC-8matchScore is miscalibrated; ties hide ambiguityMedium
- Frequency
- 54 of 151 top-1s (36%) in the 0.80–0.99 band are wrong. Only 1.0 is reliable (279/280).
- Expected
- A high score means a confident identity match; generic queries score low; ambiguity is detectable.
- Actual
- Book "Sahih", "Tafsir" and كتاب return five arbitrary titles at 0.9. "Bokhari" finds the right al-Bukhari at only 0.6. الرازي gives 0.90/0.90/0.87/0.87/0.87, so the model cannot tell it is ambiguous (feeds AE-4).
- Root cause
- library-index.ts ~274: the length-ratio formula ignores how informative the query is; ~294 caps fuzzy scores at 0.7·(1−fuseScore).
- Fix
- Score by matched IDF mass over total query IDF, plus label coverage. Penalise queries of 1–2 common tokens. Expose an
ambiguousflag when the top two are within ε.
Verifier: the "54 wrong with score ≥0.85" figure actually counts every wrong top-1 from 0.80 to 0.99. Many of these are ranking failures (LC-3/LC-4) as much as calibration failures.
Repro
LC-9Genre coverage is thin, and two English labels are misleadingMedium
- Frequency
- 38/125 genre queries empty. "principles of jurisprudence" and "jurisprudence" map to the wrong genre.
- Expected
- tasawwuf/zuhd → 23, seerah → 24, nahw → 31, "principles of jurisprudence" → 11 (Usul al-Fiqh). Category 6 labelled e.g. "Hadith Collections".
- Actual
- Transliterations and English (tasawwuf, Sufism, التصوف, الزهد, seerah, nahw, balagha, fatawa, tajweed, rijal, "legal maxims") return
[]. "principles of jurisprudence" → 12 "Jurisprudence Principles" (actually legal maxims) at 1.0. "hadith" → category 6, labelled "Sunni Books" (كتب السنة). - Root cause
- library-index.ts ~147–164:
GENRE_ALIASEShas 4 groups. The English labels for categories 6 and 12 are wrong incategory_translations. - Fix
- One alias row per genre (transliteration + English + Arabic). Relabel 6 "Hadith Collections" and 12 "Legal Maxims and Fiqh Sciences".
Verifier: the Arabic السيرة, النحو and الفتاوى work, so the gap is transliterations and English. There is no dedicated Sufism category: map zuhd/raqa'iq to 23 (Morals and Remembrances) and say so, rather than forcing a mapping.
Repro
getCategoryName(6,'en').LC-10Arabic conjunction و is not split, so genre and title words after it never matchMedium
- Frequency
- 4/7 probed Arabic genre words.
- Expected
- Genres 20, 30, 12, 26; al-Bidaya wa al-Nihaya; Tarikh al-Tabari.
- Actual
[]for genre القضاء, المعاجم, القواعد الفقهية, الطبقات and book البداية النهاية, تاريخ الرسل الملوك.- Root cause
- library-index.ts ~127–129 strips only a leading ال; والقضاء stays one token.
- Fix
- Strip the clitics و ف ب ل before ال on both query and catalog sides.
Verifier: the impact is narrower than "titles containing و". A full title written with و works (البداية والنهاية at 1.0). Only queries that drop the conjunction, or use just the word after it, fail.
Repro
LC-11Distinct records with identical English labels are indistinguishableMedium
- Frequency
- 69 same-author identical-label groups (177 books), 64–65 with more than one indexed record.
- Expected
- Rows the model can tell apart (riwaya or edition), or a canonical one first.
- Actual
- "Muwatta Malik" → three rows identical in every field at 1.0 (the Abu Mus'ab, Shaybani and Yahya recensions, all indexed). "Talbis Iblis" and "al-Adab al-Mufrad" → two identical rows each. "Nukhbat al-Fikar" → Ibn Hajar's text tied at 1.0 with a modern study of it.
- Root cause
- English titles drop the Arabic suffix (e.g. - رواية يحيى); the tool output has no variant or edition field.
- Fix
- Return the Arabic title suffix as
variant. Mark a canonical edition, or group duplicates into one candidate with all bookIds.
Verifier: the second "Nukhbat al-Fikar" row is a modern study, not a versification as first reported; a secondary work tying the original supports the finding.
Repro
LC-12Synchronous lookup blocks the event loop up to ~630 ms, plus a strong-consistency Turbopuffer round-trip per callMedium
- Frequency
- Warm sync p90 61 ms, p99 265 ms, max 630 ms (books p90 180 ms). Execute p50 152 ms, p99 1.35 s.
- Expected
- Under 10 ms CPU; availability served from memory.
- Actual
- The Fuse fallback over 8,359 books plus JS Levenshtein runs on the server thread. Misses pay the full fuzzy cost (a made-up title took 297 ms). The availability check alone can cost ~0.6 s (genre "fiqh": 0.1 ms sync, 664 ms total).
- Root cause
- library-index.ts ~283–298;
packages/engine/src/catalog-availability.ts~55–65 (consistency: strong). - Fix
- Precomputed inverted token/trigram index. Cache indexed author/book/genre sets in memory with periodic refresh, or use eventual consistency.
Repro
LC-13English titles and Latin exonyms are not resolvableLow
- Frequency
- 23/24 stretch cases; 8/8 on re-check.
- Expected
- Ibn Rushd al-Hafid, al-Raheeq al-Makhtum, Hisn al-Muslim, Riyad al-Salihin, Ihya, Ibn Khaldun, La Tahzan.
- Actual
[]for "Averroes", "Avicenna", "The Sealed Nectar", "Fortress of the Muslim", "Gardens of the Righteous", "Revival of the Religious Sciences", "Muqaddimah Ibn Khaldun", "Don't Be Sad". The transliterated forms all resolve.- Root cause
- No English-title aliases;
BOOK_ALIASESholds only the Nawawi Forty. - Fix
- A reviewed alias table for ~100 famous works and figures. Low because the model usually transliterates before calling the tool.
Repro
LC-14DB book_versions.is_indexed is false for all 8,583 versions, so lookupCatalog's own indexed field is always falseLow
- Frequency
- 8,583/8,583 versions.
- Expected
- Flag consistent with Turbopuffer.
- Actual
populateDataCachederivesrecord.indexedfrom the stale flag. The tool masks it by overwriting it with the live Turbopuffer check (118/118 correct).- Root cause
- library-index.ts ~450–477 trusts the DB flag.
apps/cli/src/index-version.tsdoes set it on completion, so the production corpus was probably loaded by a path that skipped the update, or the flag was reset later (unverified). - Fix
- Backfill the flag from Turbopuffer, or drop the DB-derived field.
Verifier, partly reproduced: the claim that other callers (validator, composer UI) receive the false value is wrong. Only lookupCatalog reads the field, so today it is dead and misleading rather than a live bug. Fixing it alone would not make unindexed filters fail fast (SR-8); that check does not exist.
Repro
is_indexed=true (0), compared with the Turbopuffer count checks.LC-15Some English labels mix Latin and Arabic scriptLow
- Frequency
- 49 book titles and 5 author names (originally reported as 2 seen incidentally).
- Expected
- Fully transliterated English labels.
- Actual
- "al-Adhkar li al-Nawawi t Mستو", "Masa'il al-Imam Ahmad wa Ishaq ibn Rahويه", about 10 "Matbu ضمن Muallafat…", broken "-ويه" endings in authors (Ibn Zanjويه, Ibn Miskويه). They reach tool output and the UI.
- Root cause
- English label generation data.
- Fix
- Find English labels containing Arabic letters and regenerate them.
Repro
3. search
The real tools.search.execute with a production createToolsContext, against live Turbopuffer, OpenRouter embeddings and Cohere rerank-v4.0-pro. Inputs were validated with searchInputSchema first, as the AI SDK does. About 180 executions; every returned row (1,280) was checked against the effective filter and the catalog.
Coverage
| Category | n | Pass | Notes |
|---|---|---|---|
| Filter semantics (author, book, genre, death-year ranges, AND vs OR, composer intersection, unknown ids, impossible combinations) | 48 | 48 | 0 violations in 560 filtered rows. Impossible combinations rejected before any DB call in 1–12 ms. |
| Query handling (tashkeel, quotations, English, transliteration, mixed, 1 word, ~2,000 chars, emoji, injection-like) | 38 | 35 | All 3 failures are unscoped hadith quotations whose primary collection is missing (SR-4). |
| Spelling-variant pairs (BM25 top-24 overlap) | 7 | 3 | Ta marbuta and hamza variants share 0/24 lexical results (SR-6). |
| Session behaviour (memo, exclusion, history, parallel, failure eviction) | 11 | 8 | Earlier-turn passages re-delivered; parallel calls lose slots; OR memo key. |
| Result quality (22 realistic questions, top 20 judged by hand) | 22 | 19 | Precision@20 ≈ 0.92. Failures are duplicate-heavy results. |
| Malformed / edge inputs the model might send | 16 | 11 | null for optional args, uppercase UUID, punctuation-only queries. |
| Latency and payload (sequential) | 15 | 15 | All under 2.1 s, exactly 3 network calls each. |
Metrics
| Rows checked vs filters/catalog | 1,280 rows, 0 violations |
|---|---|
| Latency p50 / p90 / max (sequential) | 1,134 / 1,567 / 2,078 ms. Rerank ≈ 46%, embed ≈ 32%, Turbopuffer ≈ 22% |
| Payload | Turbopuffer request ≈ 38.5 KB + ≈1.5 KB per earlier call (NotIn list); model-facing output mean 49.5 KB, max 66 KB |
| Precision@20 (manual, 22 questions) | 0.92; mean 15.4 distinct books in top 20 |
| Near-duplicate pairs (Jaccard ≥0.5) | 34 across 22 lists (9 exact repeats within one version) |
| Primary-source rank, "إنما الأعمال بالنيات…" unscoped | Sahih al-Bukhari: vector rank 91, BM25 >300. Rank 1 when scoped to the book. |
| Chunks with null author death year | 1,847,099 / 7,455,835 (24.8%) |
| Wasted slots on a related follow-up | 10/20 re-delivered from the previous turn |
| Parallel-call slot loss | 36/40 delivered for two overlapping calls in one step |
Confirmed failures
SR-1Search accepts invented author IDs (real IDs of the wrong author) with no provenance checkHigh
- Frequency
- Bench: 6/63 ID-filtered searches (9.5%) in 4 Arabic cases, 0 in English. Adversarial: 1/78.
- Expected
- Filter IDs come from lookup_catalog output (the prompt says "Never invent IDs"), or search rejects IDs never resolved in this conversation.
- Actual
- For al-Ghazali the model used authorIds [505] (al-Mu'ammal ibn Ihab); for al-Qushayri [465] (al-Hazimi). 505 and 465 are those scholars' death years. Also [102] al-Kasani, [262] al-Firyabi ×2, [580] al-Mardawi, [16] Husayn Nassar. All accepted; each returned 10–20 passages from the wrong author, and the intended restriction never applied. In old-classics-ar the Ihya and Qushayriyya were never actually searched, so the answer degraded silently.
- Root cause
validateCatalogIds(library-index.ts ~868–891) only checks existence; the session does not track resolved IDs. Model-facing passages omit author_id, so the model has no real ID to copy, and it appears to confusedeathYearAHwith an author ID.- Fix
- Record lookup_catalog outputs, composer scope and IDs from retrieved rows in
RetrievalSession. Reject unseen filter IDs with "resolve X with lookup_catalog", or echo the resolved author names in the search output so the model can notice the mismatch.
Repro
search {query:…, filters:{authorIds:[505]}} returns 10 rows, all by المؤمل بن إيهاب; [465] returns 20 rows, all by الحازمي. Only nonexistent IDs ([99999999]) are rejected. E2e: cd apps/cli && bun run bench run --agent <v28-pinned agent> --judge none --case qaf-old-classics-ar (also qaf-calendar-ar, qaf-change-genre-ar, qaf-negative-filter-ar).SR-2Follow-up turns re-deliver passages already given in earlier turns; the model sees an empty or shortened resultMedium
- Frequency
- Identical repeat query: 20/20 duplicates. Related query: 10–11/20. Unrelated query: 0/20. Expand: 4/4 neighbours re-delivered on re-expanding a passage expanded last turn.
- Expected
- Passages still in the 20-message window are excluded from the DB query, so follow-ups spend their 20 slots on new evidence.
- Actual
- Turbopuffer returns the same passages; projection strips them, so the model sees
[]or 9 of 20 and may read that as "no evidence". The UI and DB persist duplicate sources, and the full embed + Turbopuffer + rerank cost is spent. - Root cause
retrieval-session.tsseedHistoryfills onlysources, neverseen; exclusions cover only the current turn.docs/agents/chat-retrieval.mddocuments this as intended, so it is a design gap, not a code-vs-doc bug.- Fix
- In
seedHistory, add ids from history tool results within the window toseen(or a separate history-exclusion set). Keep projection dedupe as a safety net. At minimum, return a "no new passages" marker.
Repro
search {query:'فضل الصبر', filters:{authorIds:[54]}}; build history with prepareChatModelMessages; turn 2: same search (20/20 duplicates, model sees 0) or فضل الصبر على البلاء وثوابه (11/20 duplicates).SR-3Parallel searches in one step can't exclude each other's results, losing up to ~30% of slotsMedium
- Frequency
- 4 parallel near-duplicates: 55 unique vs 80 sequential in one run, 62 vs 80 on re-check (18–25 of 80 slots lost). 2 parallel: 36 vs 40.
- Expected
- Each call delivers up to 20 new passages. The prompts tell the model to search in parallel.
- Actual
- Fewer unique passages for the same 4 Turbopuffer and 4 rerank calls.
- Root cause
- tools.ts ~118–148 snapshots the exclusion filter at call start;
seenis updated only on completion, with no backfill. - Fix
- Rerank to ~40 and keep the first 20 unseen after a serialised dedupe-and-fill step.
Repro
Promise.all on one session, then sequentially on a fresh session.SR-4Unscoped exact hadith quotations never surface the primary collection, only commentariesMedium
- Frequency
- Re-check: 0/20 primary rows for 4 of 5 famous matns, 1/20 for the fifth. Same in adversarial e2e quotation cases.
- Expected
- Sahih al-Bukhari and other containing collections appear in the top 20, so takhrij answers cite the collection.
- Actual
- Top 20 is commentaries, Arba'in sharhs, lesson transcripts and fatwa sites. E2e answers name al-Tirmidhi or Muslim but cite only secondary passages.
- Root cause
packages/engine/src/search.ts: 24 candidates per branch; commentaries quoting the matn dominate both branches; no source-type diversity. Neither the tool description nor the prompt tells the model to scope to the collection for attribution.- Fix
- Prompt/tool description: for attribution, resolve the collection and search with bookIds (verified: scoped to Bukhari, the hadith is rank 1). Optionally add a category-6 filtered branch for quote-like queries in the same multiQuery and reserve slots.
Repro
search with إنما الأعمال بالنيات وإنما لكل امرئ ما نوى, من حسن إسلام المرء تركه ما لا يعنيه, لا يؤمن أحدكم…, من صام رمضان ثم أتبعه ستا…; count rows by the eight canonical collectors.SR-5Duplicate passages from multiple indexed editions and repeated texts crowd result slotsMedium
- Frequency
- Bench: 275/2,858 delivered rows (9.6%), up to 9/20 in one call. Re-check: 56/160 rows (35%) across 8 queries; book-scoped searches up to 65%, generic unscoped 0–15%. 20/394 books have 2–4 indexed versions. Expand: 5/8 windows had a neighbour ≥40% duplicated from another edition.
- Expected
- One copy per text per call, one edition per book.
- Actual
- "علاج العشق" scoped to Ibn al-Qayyim returns the same paragraph from two Zad editions and al-Tibb al-Nabawi (8/20). Bukhari-scoped "إنما الأعمال بالنيات" 13/20 duplicates across four versions. The model may treat two editions as corroborating sources.
- Root cause
search.tsanddeduplicateEvidencededupe by chunk id only; all versions are indexed into one namespace.- Fix
- Rerank to ~30, collapse rows with shingle Jaccard ≥0.8 or the same content hash, then slice to 20; apply the same to expand neighbours. One-version-per-book alone will not fix it: al-Tibb al-Nabawi is a separate catalog book excerpted from Zad.
Repro
search {query:'علاج العشق', filters:{authorIds:[14]}}; {query:'فضل صيام ست من شوال'} (ranks 1–2 byte-identical); duplicate counting with 3-word shingles, diacritics stripped.SR-6Hamza/alef and ta marbuta spelling variants give disjoint lexical resultsMedium
- Frequency
- 6/7 pairs have BM25 top-24 overlap of 0–9/24; final top-20 overlap 8–18/20.
- Expected
- Common user spellings hit the same lexical evidence.
- Actual
- Disjoint BM25 candidates. Worse, the spellings without hamza or with ه mostly hit manuscript-catalogue and colophon records stored in normalised spelling, not prose on the topic. The vector branch and reranker only partly compensate.
- Root cause
search.tskeywordQuerystrips only harakat; no أإآ→ا, ة→ه, ى→ي normalisation at index or query time. Alif maqsura on its own does little damage (21/24).- Fix
- A normalised FTS attribute at ingestion plus the same function on the query. Interim: run BM25 for both the raw and the normalised query.
Repro
SR-7Index/catalog book_id integrity errors: chunks under another book's id, or under an id missing from the catalogMedium
- Frequency
- Wider than first reported: 2 orphan book_ids (1,381 chunks, both Musnad al-Shafii) and 51 versions (42,221 chunks) stored under a sibling edition's book_id.
- Expected
- Chunk book_id equals the catalog owner of the version.
- Actual
- Musnad al-Shafii versions 21495 and 9615 sit under ids not in Postgres: shown as "Unknown Book" and unreachable by a bookIds filter, so ~78% of that book's indexed text cannot be filtered to. Tafsir al-Baghawi version 12217 sits under another id, so lookup reports the catalog book as unindexed. Version 1376 sits under Sahih al-Bukhari's id, which therefore holds 4 versions. Others include Tafsir al-Thaalabi, Ahkam al-Quran lil-Jassas and Bulugh al-Maram. 51 editions wrongly show as unindexed, and citations carry the wrong edition.
- Root cause
- Ingestion data re-keyed or stale;
format-chunk.tsfalls back to "Unknown Book". - Fix
- Re-key affected versions. Add an integrity job listing Turbopuffer book_ids missing from the catalog or mismatched with the version owner. Resolve the book name via version_id in
formatChunk.
Repro
SR-8Filters on catalogued but unindexed books return a silent 0; the author-with-no-books error is misleadingLow
- Frequency
- 4/4 unindexed-book filters (model and composer); 1/1 author with no books.
- Expected
- An explicit "source not indexed" error, like the impossible-combination rejection.
- Actual
- 0 rows after ~0.5–1 s of embed + Turbopuffer. al-Bazdawi (0 catalog books) is rejected with "No catalog books can match these AND filters… Check book authorship, genre and AH death years", which is true but does not say he has no books.
- Root cause
- The validator checks the Postgres catalog only and has no indexed check. Both "unindexed" books tested actually have their chunks under a sibling id (SR-7).
- Fix
- Use cached availability to reject all-unindexed IDs; special-case authors with no books in the message. Low because lookup_catalog already reports
indexed:false.
Repro
filters.bookIds=['ab6bbb86-1411-5d7b-830d-8fef2ac5266f'], also 5dd89763… and the same as composer scope; authorIds:[1412].SR-9deathYearAH filters silently drop the ~25% of the corpus whose authors are living or undatedLow
- Frequency
- Every deathYearAH call; 24.8% of chunks have no death year.
- Expected
- "Contemporary" includes living scholars, or the tool says undated/living authors are excluded.
- Actual
{gt:1400}returns only authors with a recorded death year; al-Munajjid, Aidh al-Qarni and similar disappear.- Root cause
buildRetrievalFilteradds['author_death_year','NotEq',null]. Prompt v28 mentions it; the tool description does not, and nothing tells the model to disclose it.- Fix
- Document it in the tool description; add an include-undated/living option.
Repro
search {query:'حكم الاحتفال بالمولد النبوي', filters:{deathYearAH:{gt:1400}}} vs unfiltered.SR-10null for any optional argument fails schema validationLow
- Frequency
- 4/4 null variants rejected. How often Gemini emits null was not measured.
- Expected
- null treated as omitted.
- Actual
- "filters: expected object, received null", "filterOperator: Invalid option", etc. Each costs a step.
- Root cause
packages/contracts/src/ai/tools.tsschemas use.optional()only.- Fix
z.preprocessto strip null keys, or.nullish()normalised to undefined.
Repro
searchInputSchema.safeParse({query:'الصبر', label:'x', filters:null}); also filterOperator, deathYearAH.gte, bookIds as null.SR-11Uppercase book UUID rejected as unknownLow
- Frequency
- 1/1 (model filter and composer filter).
- Expected
- Same result as lowercase (Zad al-Ma'ad, 20 rows).
- Actual
- "Unknown book ID D179B76B-…; resolve it with lookup_catalog." The schema's
z.uuid()accepts it first. - Root cause
validateCatalogIdsdoes an exact-case Map lookup.- Fix
- Lowercase ids before validation in both
canonicalSearchFiltersand the composer path.
Repro
filters.bookIds:['D179B76B-092B-5334-AE83-7DCEFEE1341E'].SR-12Memo key treats OR-with-one-dimension and AND as different callsLow
- Frequency
- 1/1; needs the model to flip the operator on an otherwise identical query, which is rare.
- Expected
- Memo hit.
- Actual
- 3 extra network calls and 20 lower-ranked rows (new, because of seen-id exclusion, but not requested).
- Root cause
- The memo key in tools.ts ~112–118 includes the raw operator.
- Fix
- Canonicalise the operator to AND when fewer than 2 dimensions are set.
Repro
{query:'التوسل', filters:{authorIds:[14,54]}}, then the same with filterOperator:'OR'.SR-13Punctuation-only queries pass validation and return 20 content-free dotted-line chunksLow
- Frequency
- 4/4 punctuation-only queries.
- Expected
- Rejected, or empty without embed/rerank cost.
- Actual
- 20 rows of ". . . . ." chunks at full cost (~1–1.4 s). The index really holds content-free chunks, which can also surface for normal queries containing ellipses.
- Root cause
- The query schema only does
.trim().min(1); ingestion keeps punctuation-only chunks. - Fix
- Require
/[\p{L}\p{N}]/u(not letters only: digit queries such as hadith numbers are legitimate). Drop low-letter-ratio chunks at ingestion.
Repro
؟؟؟ !!! ..., *, ..., -.SR-14Memoised repeats persist full duplicate outputs, and the [search overlap] log measures memo hits, not retrieval overlapLow
- Frequency
- Scenario with 5 calls and 2 DB calls logged as 60% overlap; otherwise always 0%.
- Expected
- No repeated stored outputs; overlap metric computed from raw rows.
- Actual
- The UI stream and persisted message carry repeated copies of the same 20 passages. The model does not see them (projection reduces them to 0), so the cost is payload and storage, not model tokens.
- Root cause
retrieval-session.tsrun()returns the cached array for memo hits;index.tscomputes overlap from deduplicated outputs.- Fix
- Return
[]or{duplicateOf}for memo hits; compute overlap from raw rows.
Repro
4. expand
Window correctness checked against direct Turbopuffer reads, plus boundaries, exclusion, filter and composer inheritance, the coordinate guard, history seeding, memoisation, chaining, concurrency, multi-edition books, usefulness and live Gemini use.
Coverage
| Category | n | Pass | Notes |
|---|---|---|---|
| Middle passages vs direct read | 30 | 30 | Exactly s−2..s+2 minus s and delivered ids; byte-identical. |
| Book start and last chunk | 14 | 14 | Books from 42 to 130,456 chunks. |
| Volume boundaries | 6 | 6 | Crosses volumes contiguously. |
| Exclusion of delivered neighbours | 6 | 6 | |
| Guessed or modified coordinates refused | 20 | 20 | Including version swaps, uppercased UUIDs, other-session coordinates. |
| Memoisation / chaining / concurrency / schema | 16 | 16 | |
| Filter inheritance from filtered searches | 21 | 21 | |
| Composer-scoped sessions | 12 | 12 | |
| Large exclusion set (300 ids) | 5 | 5 | 110–185 ms. |
| Coordinates from history | 12 | 9 | Scope change (EX-1), re-delivery (SR-2), label-less input (EX-4). |
| Multi-edition books | 8 | 3 | Expand correct in all 8; 5 had duplicate text from another edition (SR-5). |
| Usefulness on realistic cases | 11 | 6 | 2 failed upstream (wrong al-Mughni from catalog; search miss). |
| Page-order integrity (full-book scan) | 81 | 79 | 96,141 chunks; EX-5. |
| Agent follow-ups with realistic history | 6 | 6 | Gemini 3.7 Flash, prompt v28. |
| Agent in-turn expand calls | 31 | 27 | 4 guessed coordinates refused (EX-2). |
Metrics
| Windows exact vs direct read | 123/124 (the mismatch is EX-1) |
|---|---|
| Latency | p50 123 ms, p90 185 ms, max 524 ms |
| Calls with ≥1 neighbour omitted as already delivered | 39/124; 6/124 returned [] for that reason alone |
| Fabricated coordinates refused | 20/20; 1 false refusal (EX-4) |
| Books with more than one indexed version | 20/394 (Bukhari 4, Tafsir Ibn Kathir 4, Abu Dawud 3) |
| Agent in-turn refusal rate (guessed seq) | 4/31 (13%) |
| Sequence gaps or duplicates | 0 in 394 books |
Confirmed failures
No high or medium findings. expand itself is correct; its failures are edge cases and inherited data problems (see also SR-2 and SR-5, which affect expand too).
EX-1Expanding a history passage outside the current composer scope silently returns []Lowwas medium
- Frequency
- Deterministic, but only when the user changes scope mid-conversation and then asks for context around an old passage.
- Expected
- An explanatory error, as search gives. Returning out-of-scope neighbours would break the user's scope, so blocking is correct; the silent empty result is the defect.
- Actual
[]with no error. It also happens for passages from unfiltered history searches, becauseseedHistoryrebuilds the stored filter with the current composer.- Root cause
seedHistoryuses the current composer; the expand path (tools.ts ~165–188) never runsvalidateRetrievalFilters.- Fix
- Check the passage's book against the current scope and throw an explanatory error.
Repro
[]. Controls with no composer or composer still on Bukhari return 4 neighbours.EX-2The model guesses neighbour sequence numbers instead of copying them; the guard rejects correctly but each guess costs a stepLow
- Frequency
- Expand lens 4/31 in-turn calls (13%); bench 2/11; adversarial 3/12. 0/9 in follow-ups with realistic history.
- Expected
- Only delivered coordinates are expanded.
- Actual
- "Expand only coordinates from a retrieved passage…", then a retry. About +1 step and 3–5 s per guess. The model also mixed version ids between editions.
- Root cause
- The model must copy three coordinates and reaches ±2–3 to read further; the window is fixed at 2 per side.
- Fix
- Accept the passage id and resolve coordinates from session sources; add a before/after direction, or document chaining in the description.
Repro
qaf-book-en (seq 916 requested between delivered 914 and 918).EX-3Empty or partial expand results give no reason; omitted delivered neighbours look like missing onesLow
- Frequency
- 6/124 returned
[]only because of exclusion; 39/124 omitted some; bench 1/11. - Expected
- The output notes omitted neighbours ("already delivered: seq 1, 2").
- Actual
- Bare
[]; the model made another expand. The description does not mention exclusion. Separately, an identical repeat expand returns the memoised result again as duplicates. - Root cause
- Expand uses the session's exclusion filter and dedupes silently.
- Fix
- Return
alreadyDeliveredmetadata, or document it in the description.
Repro
[].EX-4History passages become unexpandable if the saved search input fails the current schema; the error message is wrongLow
- Frequency
- 1/1, but almost unreachable today: stored inputs were validated by this schema when made, so it needs a future schema change.
- Expected
- Expansion with composer-only filters, or a truthful error.
- Actual
- "…Search again if a historical passage has no sequence number", though it has one.
- Root cause
seedHistoryregisters sources only whensearchInputSchema.safeParse(part.input)succeeds.- Fix
- Register sources with the composer filter, or parse filters leniently.
Repro
{query:'حديث الدين النصيحة'} with no label (or label ''); expand its 4th result.EX-5Rare chunk-order and page-label anomaliesLow
- Frequency
- 3 boundaries in 96,141 chunks scanned, plus one spliced tail.
- Expected
- Sequence order equals book order, with correct page labels.
- Actual
- al-Mughni fi al-Du'afa v5836: after the colophon come entries from the middle of the book (pages 2:851→2:2391), which is genuinely non-contiguous text. Two others are wrong page labels on continuous text (1:72→1:63, 1:208→1:205), which mislead citations. One is uncertain.
- Root cause
- Ingestion chunk sequence and page assembly.
- Fix
- An ingestion check for backward or large page jumps and for text after colophons; re-ingest affected books.
Repro
EX-6Page-level footnotes repeat across neighbouring chunks, inflating expand output and confusing markersLow
- Frequency
- Corrected: 18/72 random windows (25%); about 100% of windows in page-footnoted editions such as Bukhari v1284 and Tafsir Ibn Kathir v23604 (25–61% of footnote characters duplicated).
- Expected
- Only the footnotes a chunk references, or dedupe within a window.
- Actual
- Bukhari seqs 3268 and 3269 carry the same page-392 footnote block, which belongs to seq 3267; 3269 then has a second, different "(١)". Up to ~7k wasted characters per expand, and markers can point to the wrong note.
- Root cause
- Chunk
footnoteshold whole-page footnotes;packages/api/src/ai/format-chunk.tspasses them through, andexpandChunkdoes no dedupe. - Fix
- Filter footnotes to markers present in the chunk text, or strip page footnotes already delivered.
Repro
5. Session and filter machinery
Mocked-model streamAgenticChat scenarios using the real tools, prepareStep, tool context and projection (live Turbopuffer reads), plus property tests of the filter builder and validator against the full catalog.
Coverage
| Category | n | Pass | Notes |
|---|---|---|---|
| Mocked-model scenarios | 30 | 25 | Failed: step limit gives no answer (AE-1), stale composer id (SF-1), follow-up re-delivery (SR-2), Turbopuffer-down lookup (LC-5), parallel slot loss (SR-3). |
| Filter builder / validator property tests | 3,000 | 3,000 | 0 cases widened past composer scope; 0 validator/filter disagreements; memo keys stable under permutation. |
| Hand-written edge cases | 23 | 21 | The 2 failures are catalog issues (LC-4 Zad al-Maad, LC-9 "Sunni Books"). |
| Exclusion list scale (NotIn up to 8,000 ids) | 6 | 6 | ~750 ms flat; excluded ids never return. |
| Validator vs indexed availability | 100 | 100 | All sampled authors are indexed. |
Metrics
| Runaway turn (mock keeps searching) | 20 steps, 400 passages, 0 answer characters, finishReason tool-calls |
|---|---|
| Prompt characters in that turn | 5,576,866 total; 557,327 at step 20 (messages only) |
| Parallel vs sequential (4 near-duplicates) | 55 vs 80 unique passages |
| Model-correctable errors sent to Sentry | 6/6 execute-time errors |
Confirmed failures
SF-1A stale composer id makes every search in the turn fail, and the model can't see or remove itLowwas medium
- Frequency
- Every search in an affected turn, but rare: reopening a chat re-hydrates the composer through
resolveFilters, which drops unknown ids. Realistic triggers are a tab left open across a catalog change, old clients, or direct API and bench callers. - Expected
- Unknown composer ids dropped consistently (and the user told); search continues in the valid scope.
- Actual
- "Unknown book ID 00000000-…; resolve it with lookup_catalog." on every search, even with the model's own valid filters. The scope instruction shown to the model omits the id, so it is told to resolve an id it never saw.
- Root cause
validateCatalogIds(composer)throws (library-index.ts ~797), whileresolveFilters(~662) andcreateToolsContext(tools.ts ~218–223) silently drop unknown ids.- Fix
- Normalise composer filters once per request (drop unknown ids, notify the client); throw only for model-supplied ids.
Repro
{bookIds:['00000000-0000-4000-8000-000000000000'], authorIds:[54]}; search {query:'التوحيد'}. Also unknown authorId 99999999 and categoryId 99999.SF-2Model-correctable tool errors go to Sentry as exceptions; the impossible-filter message claims a user scope that doesn't existLow
- Frequency
- Every execute-time validation error and every expand guard refusal. Input-schema failures are not captured (they never reach execute). 2/2 impossible-filter messages without a composer.
- Expected
- Validation failures are not Sentry exceptions; the message reflects whether a scope exists.
- Actual
captureExceptionwitherrorType: tool_call_failed. Message: "…within the user-selected source scope… user-selected scope: {} ([])".- Root cause
index.tsonToolExecutionEnd(~113–127); library-index.ts ~837–843.- Fix
- A typed
RecoverableToolErrorthat Sentry skips or logs at info; omit the scope clause when the composer is empty.
Repro
search {filters:{authorIds:[54], bookIds:['f66e9fd8-0e7a-5185-857b-020567db2646']}} with Sentry mocked.SF-3Docs out of date: benchmarks.md lists a retrieval-call budget, plan.md points to missing notes, bench checks keep a dead error stringLow
- Frequency
- 3 statements.
- Expected
- Docs match the code (only
isStepCount(20)), and references exist. - Actual
docs/agents/benchmarks.md:37says "retrieval-call budget", contradictingchat-retrieval.md:15.plan.md:3says release and managed-prompt notes are in chat-retrieval.md; they are not.apps/cli/src/bench/checks.ts:312special-cases "Retrieval budget exhausted.", which nothing produces any more.- Root cause
- Docs not updated with commit 048d2977 (budget removal).
- Fix
- Remove the budget mention and the dead string; add a prompt-update and label-promotion checklist to chat-retrieval.md.
Verifier: benchmarks.md:39 (qaf-v2 runs the production label) accurately describes the code. The real issue there is a bench-config risk, covered by PR-1.
6. End-to-end agent
Two lenses. (a) The 70-case curated benchmark (92 turns) with the current harness and prompt v28 pinned, plus a 28-case comparison against the production prompt v24. (b) 35 new adversarial cases (17 Arabic, 18 English) with v28, read by hand. Both ran Gemini 3.7 Flash through Google AI Studio, because the production gateway route was unavailable locally (see limitations).
Coverage
| Lens | Category | n | Pass | Notes |
|---|---|---|---|---|
| Bench | Completion / stability | 70 | 70 | 0 step-cap hits, max 8 steps, 0 empty answers. |
| Bench | Catalog resolution and ambiguity | 18 | 14 | All 4 ambiguity cases resolved silently (AE-4). |
| Bench | Genre filtering | 4 | 0 | 0 genre lookups, 0 categoryIds filters (AE-3). |
| Bench | Chronology | 12 | 8 | AE-7, AE-8. |
| Bench | Filter-ID provenance | 63 | 57 | 6 invented IDs (SR-1). |
| Bench | Expand calls | 11 | 8 | EX-2, EX-3. |
| Bench | Citation resolution (turns) | 92 | 86 | AE-2. |
| Bench | Insufficient evidence / grounding | 8 | 4 | AE-9. |
| Bench | Paired comparison vs v24 | 28 | 28 | v28: 24% fewer searches, same cost. |
| Adv. | Unusual author forms | 6 | 6 | Passed only because Gemini rewrote names into canonical Arabic first. |
| Adv. | Books (commentary / matn / nonexistent) | 4 | 3 | Fath al-Bari not found (LC-4). |
| Adv. | Ambiguous name | 2 | 1 | AE-6. |
| Adv. | Two-author comparison | 2 | 2 | |
| Adv. | Date filters (CE, Hijri century) | 4 | 4 | |
| Adv. | Exact quote with tashkeel / verse | 4 | 4 | Primary collections absent (SR-4). |
| Adv. | Expand (long hadith) | 2 | 2 | Ka'b ibn Malik stitched from 5 chained expands. |
| Adv. | Composer scope conflicts | 4 | 4 | |
| Adv. | Follow-up needing new evidence | 2 | 2 | |
| Adv. | Very broad question (step pressure) | 2 | 2 | Self-scoped within 2–4 steps. |
| Adv. | English question, Arabic-only concept | 1 | 1 | |
| Adv. | Author not in corpus / not indexed | 2 | 1 | AE-5. |
| Adv. | Every citation id resolves | 35 | 23 | AE-2. |
| Adv. | Direct lookup_catalog probes | 102 | 81 | See section 2. |
| Adv. | Direct search/expand/session probes | 24 | 15 | See sections 3–4. |
Metrics
| Completion | Bench 70/70 (92 turns), adversarial 35/35; all finishReason=stop; no 429s |
|---|---|
| Steps per turn | Bench mean 2.67, max 8; adversarial mean 3.3, max 7 |
| Tool calls (bench) | lookup_catalog 43, search 151 (74 unfiltered, 63 ID-filtered, 15 deathYearAH), expand 11; OR used 0 times |
| Turn latency (bench) | p50 33.6 s, p95 57.0 s; first visible text p50 18.1 s |
| Input tokens per turn | mean 59.9k, p95 195.6k; cache-read share 41% |
| Cost per turn (computed from tokens) | with cached input at 10%: mean $0.043, p95 $0.084; 70-case run $3.96 (rerank not priced) |
| Judge means (0–4, unreviewed judge) | grounding 3.83, attribution 3.97, completeness 3.84, uncertainty 3.90 |
| Unresolved citation ids | Bench 19/1,120 (1.7%) in 6/92 turns; adversarial 13/566 in 12/35 cases |
| Cross-edition near-duplicate rows | 275/2,858 (9.6%) |
| v28 vs v24 (28 cases) | searches 57 vs 75; latency p50 34.3 s vs 38.7 s; cost ~$0.055 per turn both |
Confirmed failures
AE-2Answers cite passage ids that were never delivered, and nothing validates themHigh
- Frequency
- Adversarial 12/35 cases, 13 unique bad ids. Bench 6/92 turns, 19/1,120 ids; date-boundary-en lost 12/12. Intermittent: a live re-run of 8 trials had 0 bad ids. Roughly 2% of ids, 1–13% of turns.
- Expected
- Every citation id is a delivered passage id, or bad ids are repaired or dropped deliberately before reaching the user.
- Actual
- Of the 13 unique bad ids: 5 are book_ids copied from a delivered row (none were invented from nothing), 3 are truncated UUIDs, the rest have 1–2 segments altered or spliced. One bench turn cited
book_id:seqpairs, which crashed the grounding judge. In the web UI (apps/web/src/components/chat/custom-components.tsx~52–62) unresolved ids are silently dropped, so users see fewer citation pills or none: date-boundary-en showed no citations at all. - Root cause
citation-output-middleware.tsonly repairs tag syntax. Rows expose bothidandbook_id, and prompt v28 says to cite "the result'sidsfield". Long UUIDs are error-prone to copy.- Fix
- Validate ids against session-delivered and history ids; map near-misses (book_id + seq, prefix/suffix) and drop the rest. Better: short ordinal handles mapped server-side. Fix the prompt to
id. Add a bench check that fails on unresolved ids.
Repro
cd apps/cli && bun run bench run --agent <v28-pinned agent> --judge none --case qaf-date-boundary-en.AE-3The agent never uses genre lookup or categoryIds; genre requests run unfiltered with book names in the queryHigh
- Frequency
- 4/4 genre cases; 0 genre lookups and 0 categoryIds filters in 61 lookups and 283 searches across all v28 and v24 runs. A live re-run (2 trials) did the same.
- Expected
lookup_catalog {type:'genre', query:'التفسير'}→ 3 (score 1.0), thencategoryIds:[3]or resolved mufassir authorIds.- Actual
- Unfiltered searches such as "تفسير الطبري إياك نعبد…". Only 2/7 searches naming a mufassir returned that tafsir (al-Tabari 0/2, al-Razi 0/2), though both books are indexed. Results were dominated by magazines, lessons and later works; source_scope scored 0.
- Root cause
- Prompt v28 mentions genres only in passing; nothing tells the model that "in tafsir works" is a genre restriction. Small catalog gap too: كتب التفسير returns
[]. - Fix
- A prompt rule and example (genre phrase →
type=genre→ categoryIds). Optionally warn when the query text names a catalog author or book and no filter is set.
Repro
cd apps/cli && bun run bench run --agent <v28-pinned agent> --judge none --case qaf-genre-en (also genre-ar, change-genre-ar/en).AE-1A turn still calling tools at step 20 ends with no answer text ("Interrupted")Mediumwas high
- Frequency
- 1/1 runaway mock scenario (and 1/1 with preambles: only preambles shown). The real model never got near the limit: 0/129 live turns, max 8 steps.
- Expected
- A grounded answer or closing message before the step limit, e.g. tools disabled on the final step.
- Actual
- 20 steps,
finishReason: tool-calls, empty text, 400 passages delivered, prompt grew from 401 to 557,327 characters (5.58M across the turn).classifyChatTurnmarks it interrupted; the web shows "Interrupted" with a Retry that replays everything. - Root cause
index.tssetsstopWhen: isStepCount(20);prepareRetrievalStepnever setsactiveToolsortoolChoice: 'none'. Prompt v28 tells the model the system will disable tools after eight calls, which no longer happens (PR-2).- Fix
- On the last step return
activeTools: []/toolChoice: 'none'; add an e2e test asserting non-empty text at the limit (TS-1). Cheap.
Repro
streamAgenticChat where the model emits search {query:'الصلاة N'} every step.AE-4Ambiguous titles resolved silently (al-Risala → al-Qushayriyya)Medium
- Frequency
- Silent resolution in 4/4 saved and 4/4 live runs. Strict failures: the ambiguous-book cases (2/4 saved, 2/2 live), whose notes say "Do not silently select". The ambiguous-author cases allow "resolve or disclose", so choosing Fakhr al-Din al-Razi is acceptable.
- Expected
- Disclose the ambiguity or ask.
- Actual
- "al-Risala" → al-Qushayriyya, with the EN answer claiming it is "commonly known simply as al-Risala"; the usual referent is al-Shafi'i's.
- Root cause
- The catalog hides the main candidate: bare الرسالة returns five titles tied at 0.9 without al-Risala lil-Shafi'i, and الرسالة الشافعي returns
[]. Score ties (LC-8), plus a prompt line ("resolve them using the request's context or ask") read as permission to pick. - Fix
- Prompt: disclose when top candidates are within ~0.05 and context doesn't decide. Fix catalog ranking (LC-1, LC-3, LC-8).
Repro
--case qaf-ambiguous-book-ar (also qaf-ambiguous-author-en) with the v28-pinned agent.AE-5Requested work not in the corpus (Ibn Arabi's al-Futuhat): restriction silently dropped, secondary sources presented as the FutuhatMedium
- Frequency
- 1/2 not-in-corpus cases (single saved trace).
- Expected
- Say that al-Futuhat is not in the library and label secondary summaries as such (a v28 rule).
- Actual
- Five lookups found nothing. The answer opens "In his magnum opus, al-Futuhat al-Makkiyya…" with a section of "Key Statements… in al-Futuhat", all cited to encyclopedias, refutations and magazines. Citation pills show the secondary book names, which reduces but does not remove the misleading framing. One citation id was truncated.
- Root cause
- lookup_catalog returns a bare
[]with no not-found signal; the prompt rule is unenforced. Also a trap: the top match for ابن عربي is a different, minor Ibn al-Arabi (d.617) at 0.933. - Fix
- A structured
no_match/not_indexedresult. Prompt: the first sentence must state unavailability; quotes attributed "as quoted by X".
Repro
adv-unindexed-en (asks about Ibn Arabi's al-Futuhat al-Makkiyya).AE-8"Older classical books" (Arabic) applied no date filter and cited modern authorsMedium
- Frequency
- No filter plus modern citations in 2/3 Arabic runs; the cutoff was disclosed in 0/3. The English case was correct.
- Expected
- Choose and disclose a defensible AH cutoff via deathYearAH.
- Actual
- 6 searches without deathYearAH (2 with invented IDs, SR-1); cited authors who died in 1083, 1428 and 1440 AH. One re-run used a silent
{lte:800}. - Root cause
- The Arabic phrase is treated as a style hint; no tool or prompt nudge.
- Fix
- An Arabic prompt example; a bench check that cited death years respect the disclosed cutoff.
Repro
--case qaf-old-classics-ar with the v28-pinned agent.AE-6Ambiguous-name answer adds an uncited section from general knowledgeLowwas medium
- Frequency
- 2/4 runs of the case.
- Expected
- Each reading cited or marked unavailable.
- Actual
- For "ماذا قال ابن العربي في تفسير قوله تعالى: ليس كمثله شيء؟", the Abu Bakr Ibn al-Arabi section is cited, but the "if Muhyi al-Din is meant" section describes Fusus al-Hikam positions with no citation. Other runs cited secondary critiques, which is acceptable.
- Root cause
- Model behaviour against v28 grounding rules. The catalog does signal absence (
[]for محيي الدين بن عربي); the model writes from memory anyway. - Fix
- A prompt rule for ambiguous alternatives; a judge criterion flagging long uncited paragraphs. Low: the content is broadly accurate and framed as conditional.
AE-7Hijri boundary off by one, and a CE conversion falsely claimed to be conservativeLowwas medium
- Frequency
- v28: 3 of 8 explicit-boundary cases. "before 300 AH" became
lte 300in 3/3 English runs (Arabic usedlt 300, correct). "before 1200 CE" becamelte 596in 5/6 runs. v24 also failed 3 cases. - Expected
- "before 300 AH" →
{lt:300}; "before 1200 CE" →{lte:595}with disclosure. - Actual
lte 596, while the answer claims this guarantees no author who died after 1200 CE, and the same answer notes that 596 AH runs Oct 1199 to Sep 1200.- Root cause
- Model arithmetic and inclusive-boundary errors; the prompt doesn't map before/after to lt/gt.
- Fix
- Prompt rules for lt/gt/gte/lte and CE→AH conversion; have the tool echo "Applied: death year < N AH". Low exposure: no run retrieved a violating row, and only one catalog author sits on each boundary.
Repro
--case qaf-unknown-dates-en, qaf-calendar-en, qaf-calendar-ar.AE-9History-window answers deny a message outside the window; absence overclaimed for an invented quoteLow
- Frequency
- History-window 6/6 runs; invented-quote 2/2 reruns.
- Expected
- "The start of the conversation is not visible to me"; "not found in the searched corpus".
- Actual
- "No code word was mentioned in this conversation." For the quote: "There is no source for this quote in any classical Islamic literature". The reasoning there (anachronism) is sound, so it is misleading in method rather than outcome.
- Root cause
packages/api/src/ai/agentic/history.tsslices to the last 20 messages and adds no truncation note.- Fix
- Add a system note when history is truncated; stronger abstention wording.
7. Prompt alignment
Langfuse platform/main-gemini versions 24, 28 and 29 compared with the tool schemas and code (13 checks, 4 pass). Correct in v28: no search mode, the expand rules, history coordinates and author-filter semantics.
| Load path | Label | Version loaded | Written for these tools? |
|---|---|---|---|
| Production server | production | v24 | No (old search {mode} tools, no lookup_catalog) |
| Dev server (NODE_ENV ≠ production) | latest | v29 | No (v24 derivative) |
| Bench qaf-v1 / qaf-v2 | production | v24 | No |
| Admin playground (manual pick) | retrieval-v2 | v28 | Yes, but stale (PR-2) |
Confirmed failures
PR-1No default path loads the prompt written for these tools: prod and bench load v24, dev loads v29High
- Frequency
- 3/3 default load paths.
- Expected
- Production and the qaf-v2 bench run the retrieval-v2 prompt (v28 or a corrected successor).
- Actual
- v28 carries only the
retrieval-v2label and nothing in code pins it. v24 still tells the model to retry with "semantic vs keyword search"; the schema acceptsmode:'keyword'and drops it without error. The new tool descriptions partly compensate (they say to use lookup_catalog first). Bench results for this branch measure the new tools under v24 unless the version is pinned. - Root cause
packages/api/src/ai/agentic/index.ts:67callsgetPrompt('platform/main-gemini')with no options;langfuse-client.ts~76–91 picks production/latest.apps/cli/src/bench/qaf.ts~181–185 usesproductionunlesspromptVersionis set, andagents/qaf-v2.tssets none.- Fix
- Pin
promptVersioninagents/qaf-v2.ts; move the production label in the same release as the code (or request labelretrieval-v2); log the loaded prompt version.
Repro
platform/main-gemini by label production (v24), latest (v29) and retrieval-v2 (v28).PR-2Prompt v28 contradicts the code: "kind", the removed 8-call budget and tool disabling, "omitted to fit context", the "ids" fieldMedium
- Frequency
- Static; every v28 turn (lines 27, 84, 86, 87, 91, 92).
- Expected
- Prompt matches
tools.tsand contracts:type, no budget or cutoffs, row fieldid. - Actual
- Line 86 "Use lookup_catalog with kind":
{kind:'author',…}fails "type: Invalid option" (in the live bench Gemini followed the schema, 0 schema errors). Line 92 describes "at most eight content-retrieval calls", tools being disabled and a stop after two no-progress calls; none exists. Lines 84/91 mention a shared or exhausted budget. Line 91 "some passages may be omitted to fit the model context" is mostly false. Line 27 "cite using the result'sidsfield" is wrong and contributes to AE-2. Line 87 leaves out filterOperator OR. - Root cause
- v28 was written before commit 048d2977 and the kind→type rename; plan.md lists the prompt update as pending.
- Fix
- Publish a corrected retrieval-v2 version:
type,id, filterOperator, the 20-step limit (answer before it), the death-year null caveat, collection scoping for attribution, the genre rule, ambiguity disclosure, before/after boundary rules. Remove budget, no-progress and omission language.
Verifier: "no matchScore/indexed/filterOperator guidance" is overstated. The tool descriptions explain them and reach the model on every call. The real problem is that the prompt contradicts the code. The link to AE-1 is plausible but not confirmed.
PR-3Dev default prompt v29 never names lookup_catalog or filterOperator and still mentions an exhausted budgetLowwas medium
- Frequency
- 1 stale statement, 2 omissions.
- Expected
- The dev prompt describes the current tools.
- Actual
- v29 = v24 minus "semantic vs keyword", plus 5 filter lines and "An empty result may mean… an exhausted budget". The omissions are covered by the tool descriptions, and "dimensions are ANDed" is the correct default.
- Root cause
- Label
latestmoved to a v24 derivative instead of the retrieval-v2 line. - Fix
- Same as PR-1: pin or relabel retrieval-v2.
8. Test suite
| Package | Tests | Pass | Typecheck |
|---|---|---|---|
packages/api | 485 | 485 | pass |
packages/engine | 35 | 35 | pass |
packages/contracts | 75 | 75 | pass |
apps/cli | 37 | 37 | pass |
No skips and no environment failures. The suite covers memo eviction, expand provenance, projection prefix stability and deduplication well. What it misses is listed below.
Confirmed failures
TS-1No test catches the step-limit no-answer case; the existing step-limit test encodes it as passingLow
- Frequency
- 1 test; no production-wiring e2e tests.
- Expected
- A
streamAgenticChat+ mock-model test asserting answer text at the step limit. - Actual
retrieval.test.ts(~388–499) uses fixture tools and plainstreamText, asserts 20 steps, 29 DB calls and no errors, and never checks for text.streamAgenticChatis only ever stubbed in tests. Stale composer ids and lookup with Turbopuffer down are untested.- Root cause
- Test design.
- Fix
- Add
streamAgenticChattests withMockLanguageModelV4for step-limit synthesis, stale composer id and Turbopuffer-down lookup.
TS-2Bench checks score correct behaviour as failuresLow
- Frequency
- source_scope 0 on empty-intersection en and ar (2/2 trials); expansion_provenance 0 on expansion-en.
- Expected
- Abstention cases pass when no out-of-scope rows are cited; citing an already-delivered neighbour satisfies
requireExpansion. - Actual
- Empty-intersection scope (Ibn al-Qayyim AND died before 500 AH) fails any run that retrieves anything, though the agent correctly explained the request is impossible. expansion-en turn 2 cited the next passages, already delivered in turn 1, and scored 0.
- Root cause
apps/cli/src/bench/checks.tsscopeCheckandexpansionCheck.- Fix
- Judge cited rows for abstention scope; accept already-delivered neighbours.
Verifier, partly: the third sub-claim (tool_errors failing on recovered guard errors, 68/70) is a legitimate signal of wasted steps, not a false failure.
TS-3Built-in qaf-v1/qaf-v2 bench agents can't run locally; a reused output path crashes with EEXISTLow
- Frequency
- 1/1.
- Expected
- Runs against production Gemini as documented.
- Actual
- "Your AI Gateway has authentication active, but you didn't provide a valid apiKey": the local
CLOUDFLARE_WORKERS_AI_API_KEYis a placeholder. Re-using--outputgives a rawEEXIST … linkerror. - Root cause
packages/engine/src/models.tsbuildsllmfrom that key.- Fix
- Document the gateway key or let bench agents take an explicit modelId; show a friendly "output exists" error. Local developer experience only.
Repro
cd apps/cli && bun run bench run --agent qaf-v2 --judge none --case qaf-author-en --output <path>, then run again with the same path.9. What works well
- Filter enforcement is exact. 0 violations in 1,280 returned rows; 0/3,000 random cases widened composer scope; validator and filter agreed in 3,000/3,000. Impossible combinations (author plus another author's book, author outside the date range, scope contradiction, Gregorian years) are rejected before any DB call, in 0–12 ms, with actionable messages.
- Expand is correct and cheap. 123/124 windows byte-identical to direct reads, correct at book start, end and volume boundaries; 20/20 fabricated coordinates refused; filter inheritance exact in 33/33 cases; p50 123 ms.
- The agent always finishes. 105/105 cases and 129/129 turns ended with
stop, max 8 steps, no runaway loops after the budget removal. v28 made 24% fewer searches than v24 at the same cost. - Exact canonical catalog names are reliable. 279/280 top-1s at score 1.0 are right; absent people, books and genres return
[](20/20) with no hallucinated candidates; theindexedflag matched Turbopuffer 118/118; ~90 death years spot-checked are right. - Arabic handling in catalog and BM25. Hamza/alef forms, ta marbuta, tashkeel and ʿ are folded in the catalog; tashkeel stripping makes BM25 identical with or without vowels; English, transliterated and mixed queries reach relevant Arabic passages at rank 1.
- Search quality is good when the query is right. Precision@20 ≈ 0.92 on 22 realistic questions, p50 1.13 s, exactly 3 network calls per uncached search.
- Memoisation and exclusion within a turn. Robust to label, whitespace, filter order and explicit AND; aborted calls are evicted and retry cleanly; NotIn exclusion scales to 8,000 ids with no latency growth.
- Date handling in most cases. CE→AH conversion and Hijri centuries were right in 4/4 adversarial cases; impossible date requests were explained rather than answered.
- Composer scope held in every e2e case, including deliberate conflicts, and is injected identically on every step (prefix-cache friendly).
10. Prioritised fix list
Effort: S under a day, M a few days, L a week or more, or needs re-indexing.
| # | Fix | Addresses | Effort |
|---|---|---|---|
| 1 | Publish a corrected retrieval-v2 prompt (type, id, no budget language, genre rule, collection scoping for attribution, ambiguity disclosure, lt/gt boundary rules), pin it in the qaf-v2 bench and move the production label with the release. | PR-1, PR-2, PR-3, AE-3, AE-4, AE-7, SR-4 | S |
| 2 | Catalog normalisation quick wins: strip in-word apostrophes, split the Arabic و/ف/ب/ل clitics, treat multi-word honorifics as phrases, drop 1-character tokens. | LC-2, LC-10, LC-7 | S |
| 3 | Validate citation ids server-side: map book_id + seq and prefix near-misses, drop the rest, or switch to short ordinal handles. Add a bench check. | AE-2 | M |
| 4 | Filter-ID provenance: reject author/book IDs not seen in lookup output, composer scope or retrieved rows; echo resolved names in search output. | SR-1, AE-8 | M |
| 5 | Catalog scoring: IDF-weighted partial token match, Latin transliteration folding, always merge good fuzzy candidates, an ambiguous flag, and match author names for book queries. | LC-1, LC-6, LC-8, AE-4 | M |
| 6 | Prominence prior (indexed volume per author) as tiebreak and score term, plus reviewed alias tables for top authors, genres and English titles; relabel categories 6 and 12. | LC-3, LC-9, LC-13 | M |
| 7 | Derivative ranking: penalise only titles that wrap the query, extend markers (mukhtasar, takmila, dhayl, ikmal, durar…), stop penalising originals that contain "Sharh". | LC-4 | S |
| 8 | Robustness bundle: disable tools on the final step; try/catch the availability check; normalise composer ids once per request; typed recoverable errors kept out of Sentry. Add streamAgenticChat e2e tests for each. | AE-1, LC-5, SF-1, SF-2, TS-1 | S |
| 9 | Exclude passages already delivered in the history window, and fill parallel calls after a serialised dedupe; return "already delivered" metadata. | SR-2, SR-3, EX-3, SR-14 | M |
| 10 | Index hygiene: re-key the 51 misfiled versions and 2 orphan ids, collapse near-duplicate and cross-edition rows before slicing to 20, add a normalised Arabic FTS attribute, drop content-free chunks. | SR-7, SR-5, SR-6, SR-13, SR-8 | L |
11. Method and limitations
Six independent test lenses ran against branch codex/agent-improvements @ 36803954: lookup_catalog, search, expand, machinery (session, filters, prompt, test suite), the curated benchmark, and an adversarial e2e set. Each lens used probe scripts calling the real tool code with the server environment against live Turbopuffer, embeddings and rerank. Expectations were written before running where possible. Every candidate failure then went to a separate verification pass that re-ran it with its own probe script and could correct the frequency, severity or root cause. All 53 claims reproduced, 10 of them only partly; this report uses the corrected severities and folds in the corrections. All work was read-only: no repo edits and no writes to the database, Turbopuffer or Langfuse.
- Production model route blocked. The Cloudflare AI Gateway returned 401 locally (placeholder key). All agent runs used the same model, gemini-3.7-flash, through Google AI Studio. Caching, latency and cost may differ from production Vertex; costs here are computed from token counts at $0.75/$3.75 per M, and rerank is not priced.
- Single samples. Each benchmark and adversarial case ran once (a few were re-run during verification). Model-behaviour frequencies are small-n. The v24 comparison covered 28 cases with deterministic checks only. Adversarial verdicts are manual readings; the bench judge is unreviewed.
- Not measured. Which query forms the live model actually sends to lookup_catalog (impact estimates use prompt-suggested forms and typical transliterations); how often Gemini emits null arguments; real prefix-cache hit rates; whether quota is refunded for interrupted turns; 429 behaviour of the availability check.
- Not runnable. No catalogued-but-unindexed author could be found (~220 sampled were all indexed), so that case was tested with books only.
- Harness artifact retracted. Early expand-lens agent runs showed all follow-up expands refused. The cause was a history built from
response.messages, which carries no tool messages in this SDK build. With realistic history built throughprepareChatModelMessages, 6/6 follow-ups succeeded. No repo code builds history that way. - Count caveats. Tester-level corrections: the search lens's F31 (Ibn al-Qayyim does have a Tafsir-genre book) and I12 (negative AH years for pre-Islamic poets are real data) were the testers' wrong expectations, not product failures. Some per-tool totals come from coverage tables whose units differ.
Appendix: findings that did not reproduce
None. Every confirmed claim reproduced at least in part. Partial reproductions (LC-4 Siyar, LC-14 impact, SR-14 model visibility, EX-4 reachability, EX-5 page labels, SF-1 frequency, AE-4 case notes, AE-9 invented quote, PR-3, TS-2 tool_errors) are corrected in place above.