AI Academy / Simulation

Build and evaluate a tiny retriever

Rank controlled documents by token overlap and measure retrieval accuracy. A real local experiment, not an LLM call.

You will learn to

  • Trace tokenization and ranking
  • Measure hit-at-one
  • Recognize lexical retrieval failure

Before you start

Python or JavaScript basics

Retrieve before generating

This small retriever lowercases words and ranks documents by unique query-term overlap. Everything is computed in your browser. There is no language model, provider key, training run or hidden precomputed response.

Evaluate with a fixed dataset

Hit-at-one is the fraction of labeled queries whose first result is the expected document. A small dataset can detect regressions but cannot justify broad quality claims. Add a paraphrase that shares no document words and watch the retriever fail.

From score to useful answer

Overlap misses synonyms, negation and context. A larger system might use embeddings, filters and reranking, with privacy and cost controls. The prompt comparison here checks explicit task, context, output format and uncertainty instructions. That is structural feedback, not a prediction of model quality.

Try it yourself

Enable JavaScript for this interactive activity. You can read all lesson explanations above without it.

Continue exploring