Benchmarks & System Performance

Performance & Quality Benchmarks

Standardized, reproducible performance measurements of NeuronScope across LLMs, RAG retrieval pipelines, autonomous agents, and hardware architectures.

7 / 7 benchmarks
Model / SystemParamsDevice / AdapterTelemetry ProfileHallucination RateCommit
bert-base-uncased110MNVIDIA A100 40GBFull Tensor Activations8,40098072N/A9c1a04f
gpt-4o (API Adapter)N/ACloud API GatewayPrompt/Response & Tool Spans4,200320450.3%9c1a04f
claude-3-5-sonnetN/ACloud API GatewayPrompt/Response & Context Spans3,850290420.2%9c1a04f
distilbert-base66MApple M3 MaxFull Tensor Activations3,20061041N/A9c1a04f
llama-3.1-8b-instruct8BNVIDIA H100 80GBAll-Layer Activations & Traces2,84016,4008400.8%9c1a04f
mistral-7b-v0.37BNVIDIA A100 80GBActivations + Attention Heads1,95014,8006201.1%9c1a04f
llama-3.1-70b-instruct70B8x H100 SXM5Layer Stride 4 + Tool Hooks920142,0003,2000.4%9c1a04f
Benchmark Methodology: Measurements captured under median of 5 independent evaluation runs with warm GPU cache. Token throughput includes zero-copy streaming telemetry overhead. Every row links to the sealed .nsz artifact digest and GitHub commit hash.