Idea: classify each user question into "needs a tool" vs "answerable from context" first.
The best route is rarely the cleverest route — it is the one you can evaluate.
(Demo fixture for the prototype — not a real note.)
Living AI Trace: USER → RETRIEVE → AGENT → TOOL → EVALUATE → OUTPUT
Rule: nothing important may depend on the visual — the SVG is enhancement only.
See DESIGN.md and docs/redesign/10-living-ai-trace.md
(Demo fixture for the prototype — not a real note.)
The same test set scores differently under a different rubric.
Lesson: never report just the pass count — report the rubric too.
See EXP 003 in the LAB for the simulation.
(Demo fixture for the prototype — not a real note.)
If the UI said "this document just fell out of the context window", users would guess less.
Built as EXP 004.
(Demo fixture for the prototype — not a real note.)