WholeRepo Mark
WholeRepo Zeno-VMLA All your base are belong to you.
Flash-Frustum V-MLA Kernel Live: 99.56% DRAM Bypass | 1.4s Cold Ingest →

All your base are belong to you.

The sovereign, zero-retention whole-codebase intelligence engine. Ingest 2M+ tokens in volatile GPU memory in 1.4 seconds—executing non-local architectural reasoning in a single forward pass with 0 tool calls.

Launch In-Memory Studio
$ curl -fsSL https://wholerepo.sh | sh
Cold Ingest Time
1.41sec
2.14M tokens in volatile HBM
DRAM Bus Bypass
99.56%
Geodesic Triton frustum culling
Time to First Token
24.2ms
Zeno-VMLA instant prefill
Persistent Disk Writes
0bytes
100% volatile ephemeral memory
⚡ THE DEVELOPER COPILOT INTERFACE

Zero Tool Calls. Just Native Stream.

Plug WholeRepo into your existing pipeline with standard OpenAI-compatible endpoints, an instant zero-dependency CLI, or native TypeScript / Node SDKs.

wholerepo-vmla / terminal.sh
# 1. Install sovereign Frustum zero-tool binary
curl -fsSL https://wholerepo.sh | sh
# 2. Ingest entire codebase and reason across all 1,420 files in 1 forward pass
wholerepo ask "Why does the SVG layout measurement skip embedded foreignObject?" \
--dir . \
--tier "portfolio" \ # Routes to A100-80GB Frustum HBM
--stream
# 3. Interactive multi-turn whole-codebase REPL with 0 persistent disk writes
wholerepo chat --ephemeral-hbm
GPU Worker: Modal A10G / A100 Active
FP8 Micro-Tile: 576.5 B/tok | Zero Persistent Storage
Simulated Output Stream
● Ready (Click ▶ Run)

Press "▶ Run Request" above to simulate instant whole-repo forward pass...

$ wholerepo ask "Why does the SVG layout measurement skip embedded foreignObject?"

Tokens Generated: 248

📊 Live Telemetry (Single Forward Pass)

Empirical Benchmarks
Total Latency
1.41s
Ingest + Pre-fill + Reasoning
Time to First Token
24.2ms
Zero vector indexing wait
DRAM Bypass Rate
99.56%
Triton Frustum Culling
Tool Calls Required
0calls
No fragile agentic looping
🔒 Volatile HBM Retention: 0.00% (Evaporated)
🔬 HARDWARE-AWARE ARCHITECTURE

Engineered for Volatile GPU Memory

WholeRepo discards persistent vector databases and agentic search loops. Four mathematical breakthroughs make real-time multi-million token reasoning physically possible.

Breakthrough 01 99.56% DRAM Bypass

Riemannian Frustum Culling

Standard attention thrashing chokes the GPU memory bus by loading 2M token KV caches to and from DRAM. Our custom Triton Riemannian kernel projects attention queries along geodesics, pruning unvisited token manifold shells in SRAM before they ever hit the memory controller.

Memory Hierarchy Routing 2.4x Speedup vs FlashAttention-3
Standard Dense Attention 100% DRAM Roundtrips
WholeRepo Riemannian Frustum 0.44% Bus Access (99.56% Bypass)
Host Context 2.14M Tokens
➔ Geodesic Frustum ➔
SRAM L2 Core Zero DRAM Stall
Latency: 1.41s total Triton JIT Compiled Kernel
Breakthrough 02 Phase-Space Invariance

Symplectic Shells

Long-horizon attention collapses as tokens drift in numerical precision. We map positional states into symplectic Hamiltonian phase-space coordinates $(q_i, p_i)$ preserving volume on invariant manifolds.

Liouville Measure: dΩ = dq ∧ dp ≡ const
Dispersion Drift: 0.00% 10M+ Token Horizon
Breakthrough 03 576.5 Bytes / Token

FP8 Micro-Quantization

Standard 16-bit KV caches consume 2,304 bytes per token. WholeRepo segments the attention latent into 128-element micro-tiles quantized in FP8 (E4M3), compressing state by 4x without accuracy loss.

Cache Footprint (1M Tokens) Commodity A10G (24GB)
FP16 Standard
2.30 GB
2,304 B/tok
WholeRepo FP8
0.57 GB
576.5 B/tok (4x Gain)
E4M3 Micro-Quant Tile:
Max In-Memory: 3.2M tokens on 1x A10G Zero Degradation
Breakthrough 04 100M+ Token Horizon

Sequential Ephemeral Flash-Merge

For codebases exceeding 50M tokens, WholeRepo streams raw repository chunks through an ephemeral rolling KV merge matrix. Chunks are unified into a sovereign latent representation and instantly vaporize from physical VRAM upon session completion.

Streaming DAG Ephemeral Pipeline Zero Disk Write Guarantee
Stream Chunks C₁ ... Cₙ Volatile Socket
➔
Flash-Merge Kernel Rolling KV State O(1) Memory Bound
➔
Vaporization 0 Bytes Left Complete Evaporation
Privacy Architecture: Strict Zero-Retention Enterprise Audit Compliant
🪙 TRANSPARENT GPU CREDIT BILLING

Interactive Credit & Pricing Engine

Pay strictly for active GPU cycles in volatile memory. No monthly recurring vector DB fees, zero storage retainers, and 100% predictable credit consumption.

3,000,000 tokens
~11.4 MB source code
💡 Estimated Savings vs Legacy Vector RAG Save $640 / mo
• Vector DB Hosting: $240/mo ➔ $0
• Indexing Latency: 45 mins ➔ 1.4s
Developer Tier Zero Retention
$20 / batch of runs

Includes 100 Frustum queries across your 3.0M token codebase.

Credits Consumed / Run: 24 credits
Est. Cold Ingestion: 1.41s
VRAM Footprint (FP8): 1.72 GB
Data Persistence: 0 Bytes (Evaporated)
Start Free with 50 Credits →
No credit card required • Instant API key in Sandbox

13.8M Token Multi-Repository Gauntlet

WholeRepo tested across Linux Kernel, Ladybird Browser, PyTorch, and Chromium ASTs.

Ladybird Browser (2.3M Tokens)
CSS Layout Bug Localization
Agentic RAG: 14 Tool Calls (Failed @ depth 3)
WholeRepo VMLA: 1 Pass (1.4s · Bit-Exact)
Linux Kernel v6.8 (8.4M Tokens)
RCU Lock Race Detection
Dense FA-3: OOM on 80GB HBM
WholeRepo Frustum: 2.8s Latency (99.56% Bypass)
TypeScript Compiler (3.1M Tokens)
Circular Type Resolution Trace
Graph DB / Neo4j: 12.4s Graph Hop Delay
WholeRepo VMLA: 1.8s Non-Local Attention