Skip to content
corpusframe
Get run 001
Pre-registered · protocol v0.1Published 2026-08-17 · no results exist yet

At the scale where video-frame corpora stop fitting in one GPU, most of the index is redundant.

Cull the near-duplicates first and the same recall costs a fraction of the GPU-hours. That is the claim. It has not been measured yet — the protocol that would settle it is published in full, before any run exists.

GPU-hours to index and serve 1B frames at recall@10 ≥ 0.95
—PENDING
vs FAISS-GPU IVF-PQ
—PENDING
Recall@10 floor
0.95fixed
Corpus sweep
1e7–1e9
Colour means one thing here:MeasuredClaimed, not measuredA system we don't own
§01

The claim, in two halves that fail separately

Each half has a stated falsification condition. If either trips, the run is published and this page changes.

A · cull

Redundancy is removable before indexing

Frame corpora sampled from video are dominated by near-duplicates — static shots, slow pans, repeated intros. Culling them at a stated similarity threshold shrinks the corpus without measurably moving retrieval quality on the surviving frames.

Falsified by

Culling at any threshold that shrinks the corpus meaningfully also drops recall@10 outside the stated tolerance.

B · index

The remainder indexes cheaper than FAISS-GPU does

On the culled corpus, reaching a fixed recall@10 takes fewer GPU-hours — across build and query together — than an IVF-PQ index built by FAISS-GPU on the same hardware and the same embeddings.

Falsified by

FAISS-GPU matches or beats total GPU-hours at equal recall@10, on the culled corpus, at any corpus size in the sweep.

§02

How it works

6 steps, run identically for every system under test. Every artefact at every step is published.

  1. 01

    Decode and embed

    Sample frames at 1 fps from each dataset, embed with the primary model, write fp16 embeddings plus a stable frame id to local NVMe. Publish the decode and embed scripts, and the SHA-256 of the resulting shards.

    Frame counts depend on the decode step, not on the published clip counts. The actual counts get reported; no estimate appears anywhere before then.

  2. 02

    Hold out queries

    Draw 100k query frames uniformly at random from the pre-cull corpus, with a fixed seed, and remove them from every index. Publish the seed.

  3. 03

    Build ground truth

    Exact fp32 flat search for the true top-100 of every query frame against the pre-cull corpus. Expensive, done once per corpus size, cached and published as an artefact so third parties can score their own systems against it.

    Ground truth against the *pre-cull* corpus is the strict reading: culling cannot be credited for making its own recall easier.

  4. 04

    Sweep the baselines

    Build each FAISS configuration, sweep nprobe (or efSearch), and record (recall@10, throughput, build time, memory) at every point. Keep the whole frontier, not the best point.

  5. 05

    Sweep corpusframe

    Same corpus, same embeddings, same query set, same hardware, same day. Sweep the culling threshold and the index parameters independently so their contributions are separable.

  6. 06

    Publish everything

    Raw result CSVs, the scripts, the exact environment, and the frontier plots — including the regions where corpusframe loses. Run reports are versioned in the site repository next to this protocol, so the diff is public.

    If a run contradicts the claim, the run gets published and the claim on this page changes. That commitment is the only thing making the pre-registration worth anything.

§03

Protocol v0.1 — proposed, not yet pinned

A protocol that names "recent CUDA" is not reproducible. Every unpinned value below renders as an open defect until the run node exists.

Hardware
NodeSingle node, no distributed servingMulti-node changes the economics and is out of scope for run 001.
GPU1× NVIDIA A100 80 GB SXM4PROPOSED — replace with the actual device.
GPU clocks—Locked with nvidia-smi -lgc; report the locked value. Unlocked clocks make throughput unreproducible.
CPU—Model and core count; pin at run time.
Host RAM—
StorageLocal NVMeEmbeddings read from local disk, never from object storage, so build time is not measuring the network.
Driver—Exact driver version.
CUDA—Exact toolkit version.
FAISS—Exact release or commit SHA, and how it was built — a conda FAISS and a source FAISS do not benchmark the same.
Power / thermal—Power cap and sustained temperature. A throttling A100 has produced more bogus benchmarks than any bad algorithm.
Embeddings & workload
Primary modelCLIP ViT-L/14 @ 224 px, 768-d image embeddings
Secondary sweepOpenCLIP ViT-B/32, 512-dConfirms the result is not an artefact of one embedding geometry.
NormalisationL2-normalised; similarity is inner product = cosine
Precisionfp16 on device, fp32 for ground truth
Inference costExcluded from all reported figuresIdentical across systems under test, so including it would only dilute the comparison.
Corpus sizes1e7 → 1e9 frames, half-decade stepsReported per step. A single scalar at one corpus size is not a result.
Query set100k frames held out before indexingHeld out before culling too, so culling cannot quietly delete the hard queries.
k10, with k = 100 reported alongside
Ground truthExact flat search, fp32, over the pre-cull corpus
Warm-up10k queries discarded before timing
Repetitions5 runs; report median and full min/max spreadNot mean ± σ. Spread is what tells you whether to believe the median.
ConcurrencySingle-tenant. Nothing else on the device.

WebVid

1 fps

Short stock-footage clips with captions. Heavily redundant within clips — the easy case for culling, and worth naming as such.

Research use; per-source terms apply

Instructional video with long static shots and talking heads. Different redundancy structure to WebVid, which is exactly why both are in the sweep.

Research use; per-source terms apply

Systems under test

5
  • FAISS-GPU · IVF-PQbaseline
  • FAISS-GPU · IVF-Flatbaseline
  • FAISS-GPU · Flatreference
  • HNSW (CPU)reference
  • corpusframesubject
§04

Results

Nothing here is measured. Targets must be committed to git before run 001 — a target written down after the numbers arrive is worth nothing.

MetricDefinitionFAISS-GPUcorpusframeTarget
recall@10Fraction of the true 10 nearest neighbours, by exact cosine similarity against a brute-force flat index over the same embeddings, that appear in the returned top 10. Averaged over the held-out query set.———
query throughputQueries per second at steady state, single node, warm index, measured at the batch size that maximises throughput for each system independently. Reported at fixed recall@10, never at a system’s own best recall.———
index build timeWall-clock from embeddings on local NVMe to a queryable index, including training, clustering, and any culling pass. Excludes embedding inference, which is identical for every system under test.———
GPU-hours / 1e6 queriesBuild cost amortised over a stated query volume, plus query cost. The only metric that maps onto a budget line, and the only one a buyer checks twice.———
resident index sizePeak device memory held by the index while serving, excluding the raw embedding matrix when the index does not need it resident.———
corpus reductionFraction of frames removed by the culling pass at the stated similarity threshold. Reported per dataset — it is a property of the footage, not of the system, and it will vary wildly between WebVid and HowTo100M.———
Kill conditions

Stated in advance, because a benchmark that cannot lose is not a benchmark. These are published in the run report whether or not they trip.

  • K1FAISS-GPU wins on total GPU-hours at equal recall@10 anywhere in the corpus-size sweep.
  • K2Culling at the threshold that produces the headline reduction drops recall@10 by more than 0.01 absolute.
  • K3The advantage exists at 1e9 frames but not at 1e7 — in which case the claim is about scale, and the page will say so instead of generalising.
  • K4Results do not reproduce within 5% on a second, differently-configured node.
  • K5The recall/throughput frontier crosses, and FAISS-GPU is ahead in the region most pipelines actually operate in.
§05

Reproduce it yourself

Ground truth is cached and published as an artefact, so a third party can score their own system against exactly the numbers this benchmark uses.

Score against our ground truthArtefacts pending run 001
# fetch the cached top-100 ground truth (R2, public)
curl -O https://artifacts.corpusframe.com/run001/gt-1e8.npy

# score any top-10 result file against it
python -m corpusframe.score \
  --truth gt-1e8.npy \
  --pred  my-system-top10.npy \
  --k 10
Rebuild the corpusScripts pending
# decode at 1 fps and embed, fp16 shards to NVMe
corpusframe embed \
  --dataset webvid --fps 1 --crop 224 \
  --model clip-vit-l-14 --out /nvme/shards

# hold out the query set with the published seed
corpusframe holdout --n 100000 --seed 20260817

The source repository is not public yet. Commands above are the interface the protocol commits to; they ship with run 001.

§06

Status & roadmap

Everything after Phase A is blocked on the same thing: run 001 existing.

Phase A

Protocol published, claim stated as unverified

Two pages, an email capture, zero JavaScript shipped. Live now.

Shipped
Phase 0

Pin the hardware, state the targets, run 001

Targets committed to git before the run, not after. The node replaces every PROPOSED row above.

Next
Phase B

Results page, cost calculator, Pareto frontier

An interactive frontier that shows the regions where corpusframe loses. A frontier better everywhere reads as fabricated.

Blocked
Phase C

Near-duplicate explorer, migrating from FAISS

A WebGL grid of real frames with a draggable similarity threshold, and a migration page showing the actual diff in your pipeline.

Planned
§07

One email when run 001 publishes

Whatever the result says. No newsletter, no drip, no third-party mailing-list provider — the address goes into Cloudflare KV and nowhere else.

Or write to [email protected]. No analytics, no third-party scripts, no cookies — nothing on this page phones anywhere.