Work · Case study · 01

cidx

A zero-config local code index for AI coding agents. Five read-only MCP tools answer “where is X defined, who uses Y” with locations rather than file dumps, and the incremental index is proven to equal a cold rebuild.

Alpha on PyPIPythontree-sitterSQLiteMCPJul – Sep 2026

Architecture

The system, at a glance.

saves, git opsfirst runpathsrowsrowslookupsanswersYour repositorysource files on disk · .gitignore honouredWatcherdebounce ~100 ms · coalesce per path01Cold indexerfirst full walk of the repository01Incremental enginehash → tree-sitter → one transaction02SQLite storeWAL · FTS5 · files / symbols / refs · one file, outside the repo04Query layerrank → fit ~700 tokens → shape, with a freshness stamp05CLIcidx index · query · check06MCP serverstdio · 5 read-only tools06AI coding agentgets locations, not file dumpssavesfirst runpathsrowsrowslookupsanswersYour repositorysource files · .gitignore honouredWatcherdebounce · coalesce01Cold indexerfirst full walk01Incremental enginehash → parse → writeSQLite storeWAL · FTS5 · outside the repo04Query layerrank → budget → shape05CLIquery · check06MCP serverstdio · 5 tools06AI coding agentlocations, not dumps
Drawn in the order the data moves. Numbers match the steps under “How it works”.

Try it

A real index, in your browser.

try

loading the sample index…

sample index · cidx indexing its own source

cidx indexed its own source, and the index was exported from SQLite to JSON (export script ↗). Each tab is one of the MCP tools and shows the text an agent would receive, including the token budget’s truncation marker and the index age. The box runs a JavaScript port of the query layer, not the Python tool itself, and leaves out the two edit-history ranking signals. Click a result to see who references it.

01 — Claims and evidence

ClaimMeasurementContextVerify
Incremental index equals a cold rebuildproperty suite + crash injectionHypothesis drives random edit, delete and rename storms; snapshots compared as multisetstests/convergence ↗
Test suite263 testsUbuntu, macOS, Windows × Python 3.11, 3.12, 3.13CI workflow ↗
Cold index of Django30.5 s3,043 files · 76,166 symbols · 206,801 references · warm file cache; 67.8–97.2 s with a cold cachedocs/verification.md ↗
Exact lookups22.6 ms · 27.0 msp95 for find_definition and find_references on Django, 60 samples eachdocs/verification.md ↗
Save to queryable0.33 s · 1.8 sisolated save, then p50 for back-to-back saves on Django; the whole-index re-resolution dominates at that sizedocs/verification.md ↗
Same index on Linux and Windowsfingerprint matcha pinned Django revision indexed on both, digests compared in CIcross-platform workflow ↗
Publishedalpha on PyPIApache-2.0; 16 dated architecture decision records; a changelog in Keep a Changelog formatpypi.org/project/cidx ↗

02 — The problem

Most coding agents find code by grepping and reading files. It works, but on a large repository it burns tokens on files that never end up mattering, and the agent still has to guess which of five matches is the definition.

cidx parses a repository into a symbol and reference index, keeps it fresh within milliseconds of a save, and answers over MCP with locations and signatures rather than file contents. Zero-config means exactly that: no API keys, no vector database, no Docker, no accounts. One SQLite file, stored outside the repository, and five tools that can only read.

03 — How it works

From a saved file to an answer

  1. A watcher debounces and coalesces file events; a cold indexer does the first full walk, honouring .gitignore through git plus a fixed skip-list of build and dependency directories.
  2. The incremental engine hashes the file, parses it with tree-sitter (error-tolerant, so half-written files still index), and replaces its rows in one SQLite transaction.
  3. References are resolved by a cascade: same-file scope, then explicit imports followed to their source, then a unique global name. Each reference carries a confidence tag: exact, import, or name-only.
  4. The store is stdlib SQLite in WAL mode with an FTS5 table for fuzzy symbol search, keyed by a hash of the repository path so nothing is ever written inside the repo.
  5. A query layer ranks results (match tier, kind, popularity, locality, recency), fits the answer into roughly 700 tokens with truncation markers and a freshness stamp, and recommends grep on a miss.
  6. Two consumers call the same library: the CLI, and an MCP server over stdio exposing search_symbols, find_definition, find_references, outline_file and repo_map.

Neither the CLI nor the MCP server contains logic of its own.

04 — Decisions

No embeddings in v1 (ADR-007)

Retrieval works on names, references and file structure. A vector index would add a model dependency, non-determinism and a setup step, and the questions agents actually ask (“where is X defined”, “who uses Y”) do not need it. Consequence: conceptual search is out of scope and is stated as such.

Stdlib SQLite, index outside the repository (ADR-003)

One database file per repository under the user’s cache directory, WAL mode so the watcher thread writes while the MCP thread reads, FTS5 for fuzzy search. No daemon, no service, and nothing written into a project that a user did not ask for.

The convergence invariant is a release gate (ADR-008)

The incremental index must equal what a cold rebuild would produce, including reference resolutions, after any sequence of edits, deletes and renames. A property-based suite proves it, a crash-injection test shows the index can never be half-written, and cidx check lets a user prove it on their own machine.

Five read-only tools over stdio (ADR-006)

Read-only by design: a hostile file in a repository has no blast radius because there is nothing it can make cidx write or execute. Stdio keeps the setup to one line in an MCP client config.

Benchmark rules written before the first run (ADR-005, ADR-013)

The evaluation harness ships in the repository with honesty rules fixed in advance: raw JSONL logs published with every result, every number recomputable from the logs alone, losses shown at the same prominence as wins, and a dev/holdout task split that is never tuned against.

05 — What went wrong

Stated on purpose.

The cold start was O(n²)

cidx serve on a repository with no index built it through the watcher’s reconciliation sweep, which recomputed whole-index resolution and spawned git check-ignore once per file. Django had not finished after 400 seconds, and a doubling test showed 3.8× time per 2× files. Batching the whole-index work to the end of the sweep brought the same cold start to about two minutes (ADR-015).

Three foreign-key columns had no index

Replacing one file’s rows scanned the whole symbols and refs tables. On Django a re-index ran past 11 minutes and every single-file save paid a 360 ms scan. Three indexes took the re-index to 46 seconds and the per-file delete from 361 ms to 1.7 ms (ADR-016).

A write-lock race that bypassed the busy timeout

Write transactions used a deferred BEGIN and read before writing, so a second connection committing in between failed at once with SQLITE_BUSY_SNAPSHOT. Seen on CI as the watcher’s consumer thread dying during its first sweep. Every mutation now begins IMMEDIATE, and an unhandled exception in the watcher thread is logged instead of killing it silently.

The convergence check had a blind spot

It compared row sets, so an index missing one of several identical rows still passed. Minified JavaScript yields distinct functions with identical rows; Django’s vendored select2 has eleven. Snapshots are now multisets, so counts must match too.

ESM import specifiers never resolved

TypeScript and ESM imports that carry an emitted extension (./core.js, as NodeNext requires) built candidates from the specifier verbatim, so on zod 0 of 53,535 references carried the import confidence. Such specifiers now map to the TypeScript source.

06 — Limitations

  • TypeScript type-level declarations (interface, type, enum) are not indexed in v1; value-level code is.
  • CommonJS exports are not tracked as bindings; require() calls do appear as references.
  • No semantic or conceptual search, by design.
  • Namespace imports (import * as core) are not followed; core.thing() resolves by unique global name at best.
  • There is no type checker. Confidence tags say how a reference was resolved; they do not promise precision.
  • After a save, freshness cost grows with the repository: the saved file’s rows land in about 0.3 s, but references are re-resolved across the whole index, about 1.5 s on Django.
  • Ranked search and the repository map scan: about 135 ms p95 on Django, over the 50 ms target at that scale.

07 — What I’d do next

  • Incremental reference resolution, so save latency stops scaling with repository size.
  • Type-level TypeScript symbols, which the limitations list makes the most-requested gap.
  • Publishing a first benchmark run under the methodology already in the repository.

Last verified 2026-09-16 · numbers reported as measured, with their context and source