Benchmarks
Benchmarks
Reproducible measurements of NeuronScope on real hardware. Every row links to the .nsz artifact and the exact commit that produced it.
8 / 8 rows
| Model | Params | Device | Probes | Commit | |||
|---|---|---|---|---|---|---|---|
| distilbert-base-uncased | 66M | A100 40GB | activations | 1,840 | 620 | 41 | 9c1a04f |
| bert-base-uncased | 110M | A100 40GB | activations | 1,210 | 980 | 72 | 9c1a04f |
| gpt2 | 124M | A100 40GB | activations+attention | 780 | 1,440 | 168 | 9c1a04f |
| bert-large-uncased | 340M | A100 40GB | activations | 420 | 2,860 | 214 | 9c1a04f |
| gpt2-medium | 355M | A100 40GB | activations+attention | 240 | 3,820 | 512 | 9c1a04f |
| distilbert-base-uncased | 66M | M3 Max | activations | 240 | 610 | 41 | 9c1a04f |
| llama-3.1-8b | 8B | H100 80GB | activations (stride=2) | 110 | 18,400 | 1,280 | 9c1a04f |
| distilbert-base-uncased | 66M | CPU (16c) | activations | 72 | 580 | 41 | 9c1a04f |
Methodology: batch size 32, tokens per row 128, ring buffer 1 GiB, warm cache. Numbers are the median of 5 runs after 100 rows of warm-up. Reproduce with bench/run.py at the commit in the last column.