Backtesting#

h5i-db-backtest simulates venues against recorded market data. It never routes a live order.

Three properties shape the design, and each is a test rather than an intention:

The tables#

Market data is venue-neutral. A Polymarket loader and a Hyperliquid loader both produce the same tables, and the kernel never learns which vendor a row came from.

Table Holds
book_deltas snapshots, incremental level changes, and explicit gaps
trades prints, with an optional aggressor
bars aggregates
instruments one row per outcome, so a categorical market is N rows
resolutions how a market ended; never read on the strategy path

Every market-data table is time-indexed on ts_init, the column replay sorts by, so a range scan prunes on exactly the right column.

Numbers are Float64 on disk and fixed point in the kernel. The conversion is exact for every value the fixed-point type can represent (nine decimal places, magnitudes below about 9e9); a test walks the entire 0.0001 tick grid to confirm it rather than assuming.

Snapshots are grouped#

A snapshot is many rows sharing an event_index, with is_last on the final one. Applying half a snapshot would leave a crossed or hollow book, so a reader that meets a truncated one refuses it instead of reconstructing something plausible from the fragment.

A run#

use h5i_db_backtest::{run_in_fork, RunSpec, Money};

let mut strategy = SignalReplay::new(intents)?;
let report = run_in_fork(
    &db,
    RunSpec::new("momentum-001", Money::from_units(10_000)?)
        .window(window)
        .read_at(ReadAt::Snapshot("2024-q1".into()))
        .minimum_coverage(0.95),
    &mut strategy,
    |engine| engine.fee_model(Box::new(PredictionMarketFees::new(0.07)?)),
).await?;

The run creates a fork, replays inside it, and writes its results there:

Table Holds
bt_run the manifest: pin, config digest, cash, how far it simulated
bt_orders every order and its final status
bt_fills every execution
bt_positions where it finished, with settlement attribution
bt_equity the equity curve

Because results are ordinary tables on a branch, everything else already works: fork_diff compares two runs at fill level, a cross-fork scan aggregates a sweep, promote publishes a blessed run, and drop_fork_tree disposes of the rest.

bt_fills is authoritative. Positions are a fold over it and nothing else, so a stored run can be rebuilt and checked rather than trusted.

From a run to a tearsheet#

from h5i_db import quant

fork = db.fork("bt-momentum-001")
series = quant.from_levels(fork, "bt_equity")
quant.tearsheet(series, path="run.html")

The statistics are the same empyrical-parity set documented in Quant workflows; nothing about them is backtest-specific.

Settlement#

Settlement is not an event in the replay stream. It is a policy applied after the run, gated on one question: did the run reach the instant the resolution became observable?

A three-day replay of a six-month market ends holding a position. Marking it to the eventual winner books a profit nobody trading those three days could have collected, so settlement applies only when simulated_through >= observable_at. Otherwise the mark-to-market result stands and the report says why.

Both numbers survive. market_exit_pnl is what the position was worth at the last mark, settled_pnl is what it became at resolution, and the difference is reported as an explicit adjustment rather than folded in silently.

Venue models#

Four small traits carry every behavioural variation:

Trait Decides
FeeModel what a fill costs
FillModel what book an order meets
LatencyModel when the venue hears about an order
VenueModule periodic processes (funding, liquidation)

New behaviour is a new implementation, never a new flag. FillModel has one escape hatch worth knowing: book_for_fill may return a synthetic book, and matching runs against it unchanged, so slippage, bar-derived quotes and synthetic depth all reuse one matching path.

PredictionMarketFees implements the curved fee these venues actually charge, rate · quantity · p · (1 - p), which peaks at even odds and vanishes at certainty. A flat notional × rate is the wrong shape and overcharges the tails, which is where these markets trade most.

Strategies#

Tier 1, signal replay is the strategy as data: a list of timestamped order intents replayed through the full matching, fee and latency path. It covers most systematic research with no callback code and no language boundary in the loop.

Tier 2 is the Strategy trait, for path-dependent logic. Its callbacks receive a Context that exposes the clock, books, positions and cash -- and nothing else. There is no route from it to a resolution.

Two orderings inside the loop are load-bearing:

What is not here#

No live order routing, no brokerage adapters, no portfolio optimisation, no plotting API. The boundary is simulation and evaluation; see ROADMAP_QUANT.md §11 for the full list and the reasoning.