An English question comes in; a cited, checkable answer goes out. In between are five stages. This page is not about what the stages are — it is about the choices made at each one, and why.
One line runs through all of it: anything code can judge is never left to the AI.
The model understands and writes. The code does the checking. Most of the decisions below trace back to that sentence.
This is the stage most likely to go wrong. Choose the wrong words and every later stage is faithfully serving a mistake — the material never comes back, so there is no answer to be had.
Pāli is heavily inflected: one word shows up in dozens of declined forms, so a literal match on the question's own wording finds nothing.
On the English side the words are chosen by the model reading your question, not by lookup in a bilingual term table. That is a deliberate choice, and it rests on something specific about English: a great many English Buddhist terms are either Pāli loanwords or stable, settled renderings, so the gap the model has to cross is short.
⚠️ It is also a constraint, stated plainly: the term vocabulary this project draws on exists only in Chinese, so the Chinese side of the site has a lookup layer that the English side does not. English leans on the model for this one stage.
mindfulness → sati; dependent origination → paṭiccasamuppāda; the five hindrances → nīvaraṇa.
Where a term has no settled English form, the Pāli is usually what an English reader has already met — which is why typing the Pāli directly works well here.
Also in this stage: spelling repair. When the database can expand no inflected form of a word at all, the correct spelling is recovered from the corpus itself — behind four separate guards, so that a word that was already right never gets "fixed" into a wrong one.
The cost structure of retrieval means the best answer to "how much should we ask for" is not "just enough".
Cross-word intersection: the question "which five faults were listed against Ānanda" is answered precisely by asking which passage holds those words together.
Top-up rounds: measured over thirty questions, twenty-one needed only one round, five took two, four ran the full three — easy questions skip the cost automatically; only hard ones pay it.
B.2 · WHICH DICTIONARY GETS IN IS A SCHOLARLY JUDGEMENT
For a single Pāli word the database can return a dozen dictionaries' worth of glosses, and the context budget is finite. Which ones get in directly determines how reliable the gloss is — and leaving that unmanaged means handing the decision to return order.
Why the cap is needed: CPD can return thirty-odd entries for a single word (kassapa does). Uncapped, the glosses drown the text.
What it was like before: within one language, which dictionary got in depended entirely on the order the database returned them in. The tier takes that decision back.
⚠️ The authoritative list is a scholarly judgement, not an engineering one — the program does not extend it on its own.
The model may use only what stage B brought back. Nothing is filled in from memory.
Proper names stay in Pāli transliteration here. On the English side that is simply what a reader expects; it also sidesteps a problem the Chinese side has to manage, where a familiar-looking traditional rendering may not point at the same person or the same sutta.
Every claim carries a coordinate back to the passage it came from, and the answer stays inside the retrieved material.
Mahāgovinda stays Mahāgovinda.
This stage cannot be made error-free: a model slips — a wrong name, a wrong sutta number, a title attached to the wrong work. So the strategy is not to stop C from erring. It is to have D catch it before you see it.
The answer is written by a model, and a model's errors are probabilistic — they never run out. Hand the checking to a model as well and you are backing one probabilistic thing with another.
Move the checking into code and the character of it changes: one bug lets an entire class of error through a hundred per cent of the time — and, read the other way round, one correct check stops an entire class a hundred per cent of the time. Determinism is the only thing worth having in a checking stage.
Six checks, each one mechanical:
Check 3 is the one you can see working. A coordinate that is present but not clickable is a coordinate this stage could not confirm — it was deliberately left unlinked rather than quietly dropped.
The Chinese pages run a seventh check, for a pair of Pāli words whose Chinese renderings get swapped for each other. It has no English counterpart — the failure it catches only exists in Chinese.
A system that is always certain and a system that always hedges are both unusable. What makes one usable is this: it has a view about how firm its own footing is, and it says so.
⚠️ And D cannot reach this: D checks whether an answer is well formed. Search the wrong thing and every check can pass — formally impeccable, answering a different question than the one you asked.
The search is narrated as it runs. Changing a root, overturning the previous round, noticing that the material does not match the question — each of those is said out loud, in plain language, as it happens.
⚠️ Stated plainly: the English pages do not yet carry an uncertainty flag. The two signals that raise one on the Chinese side both depend on features of Chinese phrasing, so they never fire here. Building it for English means devising English signals, not translating those two. Until then, treat the narration as the only signal you get — and when an answer matters, rephrase and search again.
"That name is not in our term list, so I am splitting it literally — let me try."
"Hold on — the two passages I found are about the Bodhisatta's reflection, which is not what you asked."
"Trying a different root."
Laying the doubt out is itself the evidence of reliability — what you see is us being exacting on your behalf, rather than a vague disclaimer.
What is drawn above is how one question gets answered. This part is a different matter: when we want to touch any one of those stages, what stops us from fixing one thing and quietly breaking three — with nobody able to tell.
A criterion once used here was "the answer contains these two words", and it wrongly condemned two good answers; it had to be narrowed to "does the answer open with this sentence".
We have also seen a question come out worse on a re-run with nothing changed — there is a large model in the chain, and one question flipping is not a regression. Rephrase it, test again, and only act once the decline is reproducible. Chasing noise into the code only makes things worse.
⚠️ The English pages have no regression bank of their own yet. The bank above is in Chinese. An English-side change is therefore less well guarded than a Chinese-side one — stated here rather than left for you to find out.