AI ENGINEER · EXPERIMENTS

LAB

Small AI-engineering questions, run as deterministic simulations. Everything here runs in your browser — no APIs, no keys, no benchmark claims. Break things on purpose; that is the point.

USER RETRIEVE AGENT TOOL EVALUATE OUTPUT
Diagram: the experiment trace from run to evaluation.
  1. EXP 001 Agent Trace When an agent answers wrongly, which step of the trace actually failed? AGENTS SIMULATION
  2. EXP 002 RAG Retrieval If the best chunk is corrupted, does the system still answer correctly? RAG SIMULATION
  3. EXP 003 Evaluation If you change the rubric, did the same system really get better? EVALUATION SIMULATION
  4. EXP 004 Context Window When does the important information silently fall out of the context window? CONTEXT SIMULATION