Trailhead logo

Rust · inverted index · TF-IDF

Search that shows
its own work.

Trailhead is a full-text search engine in Rust, built from scratch and small enough to read end to end. Tokenizer, inverted index, TF-IDF ranking, no black box. Here is the real thing, running in your browser.

How to use this playground
Type a query in the search box below. Trailhead tokenizes it, looks each term up in an inverted index built over a small built-in corpus, and ranks the matching documents by TF-IDF, so rare and repeated terms carry more weight. Results update as you type, each line showing the document and its score, highest first. Try a single common word like fox, then a multi-term query like search ranking to watch the scores combine.
  1. Type something above.

tokenized, indexed, and ranked entirely in this page

What it does

Four pieces turn a folder of text files into ranked search results.

Tokenizer

Lowercases text, splits on non-alphanumeric characters, and optionally filters a small stopword list.

Inverted index

Maps each term to the documents it appears in, with per-document term counts, so lookups skip everything that does not match.

TF-IDF ranking

Scores documents by term frequency times inverse document frequency, so rare and repeated terms carry more weight.

Simple CLI

Two commands, index a directory of text files, then search it with a ranked top-N list of results.

Why it looks the way it does

No external search service and no dependency on a search library. Every choice keeps the pipeline small enough to read end to end.

From scratch

The tokenizer, inverted index, and TF-IDF scorer are all written in a few hundred lines of Rust. Nothing hides behind a black-box search crate, so you can follow a query from raw text to a ranked list.

TF-IDF, not just matching

Ranking is term frequency times inverse document frequency, summed over the query. IDF is ln(N / df), so a rare term outweighs a common one and a short document outranks a long one on the same match count.

Readable JSON index

The index serializes to plain JSON with serde. It is not the most compact format, but you can open the file and see exactly which term points to which document.

Deterministic tests

The suite covers the tokenizer and the index: exact-match queries, TF-IDF ordering, absent terms, multi-term score combination, and a save and load round trip.

Index it. Search it.

Two commands. Build the index over a directory of text files, then query it with a ranked top-N list.

Index

trailhead index <dir> tokenizes every .txt and .md file under a directory and writes trailhead.index.json in the current directory.

Search

trailhead search "<query>" --top N loads the index and prints the top-N documents by TF-IDF score, highest first.

$ trailhead index ./notes
indexed 4 document(s) with 30 unique term(s), saved to trailhead.index.json
$ trailhead search "inverted index ranking" --top 5
1. ./notes/index.md  (score: 0.3961)
$ trailhead search "document frequency" --top 3
1. ./notes/search.md  (score: 0.4159)