Benchmarks & System Performance
Performance & Quality Benchmarks
Standardized, reproducible performance measurements of NeuronScope across LLMs, RAG retrieval pipelines, autonomous agents, and hardware architectures.
7 / 7 benchmarks
| Model / System | Params | Device / Adapter | Telemetry Profile | Hallucination Rate | Commit | |||
|---|---|---|---|---|---|---|---|---|
| bert-base-uncased | 110M | NVIDIA A100 40GB | Full Tensor Activations | 8,400 | 980 | 72 | N/A | 9c1a04f |
| gpt-4o (API Adapter) | N/A | Cloud API Gateway | Prompt/Response & Tool Spans | 4,200 | 320 | 45 | 0.3% | 9c1a04f |
| claude-3-5-sonnet | N/A | Cloud API Gateway | Prompt/Response & Context Spans | 3,850 | 290 | 42 | 0.2% | 9c1a04f |
| distilbert-base | 66M | Apple M3 Max | Full Tensor Activations | 3,200 | 610 | 41 | N/A | 9c1a04f |
| llama-3.1-8b-instruct | 8B | NVIDIA H100 80GB | All-Layer Activations & Traces | 2,840 | 16,400 | 840 | 0.8% | 9c1a04f |
| mistral-7b-v0.3 | 7B | NVIDIA A100 80GB | Activations + Attention Heads | 1,950 | 14,800 | 620 | 1.1% | 9c1a04f |
| llama-3.1-70b-instruct | 70B | 8x H100 SXM5 | Layer Stride 4 + Tool Hooks | 920 | 142,000 | 3,200 | 0.4% | 9c1a04f |
Benchmark Methodology: Measurements captured under median of 5 independent evaluation runs with warm GPU cache. Token throughput includes zero-copy streaming telemetry overhead. Every row links to the sealed
.nsz artifact digest and GitHub commit hash.