LAB
Small AI-engineering questions, run as deterministic simulations. Everything here runs in your browser — no APIs, no keys, no benchmark claims. Break things on purpose; that is the point.
- Agent Trace When an agent answers wrongly, which step of the trace actually failed?
- RAG Retrieval If the best chunk is corrupted, does the system still answer correctly?
- Evaluation If you change the rubric, did the same system really get better?
- Context Window When does the important information silently fall out of the context window?