NovaMart Benchmark

NovaMart

A Causally Consistent Simulated Enterprise for Measuring Tribal Knowledge Extraction

How much of a company's undocumented knowledge can an AI agent excavate from the data estate alone: the code, the git history, the warehouse, the logs, the dashboards?

Paper Code Submit

About NovaMart

AI agents are moving into the daily work of enterprise engineers and analysts. Doing that work properly requires the company's tribal knowledge: which undocumented filters go into a metric, which of several views is canonical, which half-finished migration decided whose data moved. Agents that lack it fail silently, shipping numbers that look right while disagreeing with every existing report.

NovaMart is a simulated e-commerce retailer that is executed rather than authored. It is built from two inputs. The first is real shopper behavior: the public REES46 event stream, replayed hourly as ambient traffic through a live application and database. The second is the tape: an authored script of the company's story, written by engineers reproducing failure patterns seen in production (for example: a payment provider that delivers every confirmation twice for three hours, exposing the missing idempotency check, a vendor whose weekly price uploads are silently discarded for a month, three teams that each use their own definition of an active customer). LLM-powered engineers then work through that script against the running application, on one shared clock, for 3.5 simulated months. We never write the data ourselves: rows, logs, queries and dashboards pile up as a side effect of the work, the way they do in a real company. Then we freeze everything. Every row, log line, query and dashboard in the estate traces back to that one run.

What's inside

Application repo, full git history112 commits
Warehouse tables32
Warehouse rows878,918
Runtime + query log lines5,168,645
Redash dashboardsincluded (full export)
Simulated history3.5 months
Audited gold claims51
Baseline books from the paper9 (3 systems × 3 runs)

Every order and payment traces back to a real browsing session in the replayed REES46 stream. Every claim is re-verifiable against the shipped estate without the generator. The claims and the scoring harness ship in the GitHub repo; the frozen estate is a versioned Hugging Face dataset plus the application repo on GitHub. The README walks through local (docker compose) and GCP setup.

Set it up on GitHub →

How evaluation works

Each system gets identical read-only access to the frozen estate and a single brief: write the company's missing knowledge book. 51 audited gold claims, each with a recomputable evidence chain, decide how much it found. The claims split into narrated (20: the rubric can be satisfied from text surfaces alone, i.e. code, git history, docs) and excavated (31: requiring computation over the warehouse, logs, or query history). In our baseline runs (three systems, three runs each), the systems differed most where knowledge had to be computed rather than read: what separated them was not reading ability but excavation ability.

Each agent's book is passed to an LLM judge (the same model and prompt for every entry), which scores it against every claim's rubric: one verdict per claim. Claim recall is the fraction of the 51 claims the book satisfies. A system's headline score is mean recall over its three runs; pass³ counts the claims it solved in all three. The judge protocol, its stability analysis, and the confidence-interval recipes are in the paper. Every verdict ships with its entry under submissions/, and the leaderboard is compiled from those records (compile_leaderboard.py --check, enforced in CI), so every number on this page can be recomputed.

Score your agent's book →

Example: a gold claim, verbatim

Claim HC-02, exactly as released. This is what tribal knowledge looks like: an engineer cleans something up by hand with a single UPDATE. No code path performs that operation, and nobody beyond that engineer knows what was released or why; the knowledge sits in one person's head. The only trace is in the database query log, and the only way to learn it is to go through that log.

"On 2019-12-29 an engineer manually released fraud-held orders with a hand-run UPDATE (status 6 back to 1, for held orders under $2,600) — an operation no code path performs; only the database session records its execution. The same day, a commit raised the fraud-hold threshold from 0.70 to 0.85: the release and the threshold change were one coordinated cleanup."

narration: excavated · type: dynamic · acquisition: hard-derive · source: codebase + data warehouse · requires: query history, commit history

The evidence chain spans two sources and is recomputable against the shipped estate:

-- 1. the database query log: find the engineer session
SELECT textPayload FROM `<warehouse>.novamart_logs.db_queries`
WHERE textPayload LIKE '%engineer%' AND textPayload LIKE '%status = 6%'
ORDER BY timestamp
-- returns 2019-12-29T15:00 [engineer-backfill:maya]
--   "UPDATE orders SET status = 1, updated_at = NOW() WHERE status = 6 AND price < 2600;"
--   preceded by two survey SELECTs over held orders in the same session

-- 2. the codebase: tie it to the same-day commit
git show 1cb8721 -- novamart/constants.py
-- 2019-12-29: "Raise FRAUD_HOLD_THRESHOLD from 0.70 to 0.85 and release held orders under $2600"

A book earns this claim only if it states all of the following:

  • Held orders were mass-released by a manual SQL operation run by an engineer, not by any code path.
  • The operation is attributable to an engineer session recorded in the database query log.
  • It coincided with the fraud threshold being raised to 0.85 in late December.

and avoids the contradiction:

  • "The release of held orders was performed by an application endpoint or scheduled job."

Citation

@inproceedings{dhar2026novamart,
  title     = {NovaMart: A Causally Consistent Simulated Enterprise
               for Measuring Tribal Knowledge Extraction},
  author    = {Dhar, Susnato and Saket, Srijan and Mehrotra, Rishabh},
  booktitle = {NeurIPS 2026 Workshop on Agentic AI Benchmarks and
               Applications for Enterprise Tasks (AABA4ET)},
  year      = {2026},
  url       = {https://novamartbench.com}
}

Questions: please open an issue on the GitHub repo; or email susnatodhar10@gmail.com.