EnglishРусский Map

Neutrino-1 8B

title
Neutrino-1 8B
created
2026-07-29
tags
clippings

Available 2026-07-27

The flagship of the Neutrino family: 36 decoder layers behind a coded ternary-family container that serves a datacenter GPU, a MacBook, and a desktop CPU from one artifact.

2.56 GB

Download, lossless

1/8

The bits of fp16

72.1

MMLU

763 tok/s

Spec decode, H100

Neutrino-1 8B

Neutrino-1 8B is an 8.19B-parameter decoder-only transformer that ships as one 3.88 GB file. Every one of its 252 transformer linears is stored in a proprietary ternary-family weight format eight times smaller than fp16; the weights stay bit-packed at rest and are decoded inside the matrix kernels, so nothing in the decode path is stored as fp16 or fp32 weight material.

Small weights change the serving economics. Single-stream decode is bound by how many bytes move per token, so a 3.88 GB working set decodes at rates a 16 GB fp16 artifact cannot reach on the same memory system, and the whole model fits beside its KV cache on an 8 GB GPU or a 16 GB laptop. The same container serves every platform below without conversion.

A dense decoder-only transformer. Grouped-query attention holds the KV cache at a quarter of the query width, 288 KiB per token at the live fp32 default, so a 4k-token session costs 1.21 GB of cache beside the 3.88 GB of weights.

Base modelApache-2.0, Alibaba Cloud

Qwen3-8B

Parameters6.95B coded projection weights, 1.24B int8 embedding, 0.3M norm

8,190,735,360

Decoder layers

36

Hidden width

4,096

Feed-forward widthgated (SwiGLU), three linears per layer

12,288

Attentiongrouped-query 4:1, head width 128

32 query heads, 8 key-value heads

KV cachelive default: fp32; 1.21 GB at 4k context, 9.66 GB at 32k

288 KiB per token

Position encodingapplied across the full 128-wide head

rotary, base 1,000,000

Normalizationplus per-head query/key RMSNorm inside attention

RMSNorm, eps 1e-6

Context length

40,960 tokens

Vocabulary

151,936

Embeddingsinput embedding and output head are separate tensors

untied

Geometry as read from the shipped container's header.

Where the bytes live

Only the transformer linears carry the coded format. The two embedding tensors stay int8 because their rows are read one token at a time, not multiplied against the full activation stream, and the normalization weights are too small to be worth coding. A third of the file is vocabulary.

252 transformer linearsthe coded ternary-family lane, 67.2% of the file: query, key, value, and output projections plus the gate, up, and down feed-forward linears, 7 per layer, 72,351,744 bytes per layer

2,605 MB

Token embeddings32.1% of the file: two untied int8 tensors of 151,936 × 4,096, input embedding and output head, one scale per row

1,245 MB

Per-row metadata0.6%: row dimensions, scales, and row sums

25 MB

Normalization weights145 tensors, kept float32: four per layer plus the final norm

1.2 MB

Container header

60 bytes

Byte budget of the 3,875,404,812-byte container, by tensor class.

Inside the coded lane

Across the 6.95B coded weights, 62.63% sit at zero and the remainder splits 18.68% plus to 18.69% minus: sign-balanced to a hundredth of a point with no constraint asking for it. The balance is not uniform in depth. The gate and down feed-forward projections spike to 70 to 72% zeros in layers 1 through 3 while all four attention projections hold within about one point of 62% at every depth: the early feed-forward blocks shed weights the network does not need, and attention keeps a constant code density from layer 0 to layer 35.

gateupdownquery, key, value, output

60%65%70%08162435

60%65%70%08162435

Share of coded weights at zero, per projection, across all 36 layers.

One container, three doors. The download is a coded transport of the container, not a compressed copy of an fp16 model: it expands bit-exactly to the file every runtime executes, and that one file is what runs on a datacenter GPU and on a laptop alike.

Downloadcoded transport, 2,559,822,594 bytes; expansion is bit-exact

2.56 GB

On diskone container, 3,875,404,812 bytes

3.88 GB

Distribution surfaces

pip engine

24.9 tok/s on an Apple M5, CPU only, 9 threads

The one-command door: pip install fermion-research downloads the container and the platform-matching native binary. CPU runtimes for macOS arm64 and Linux x86-64, with a bit-exact torch reference path underneath.

GGUF pack + CUDA fork

30.7 tok/s on an NVIDIA L4, 4.68 GiB at 4k context

The llama.cpp door: the container converted to GGUF with our weight types, loaded by our public llama.cpp fork. Runs llama-completion and llama-bench, with full CUDA offload.

MLX pack

33.7 tok/s on a base M5 MacBook

The Apple-silicon door: Python-native runtime with custom Metal kernels. The container is memory-mapped and the packed planes are decoded inside the GEMV kernels.

The release battery runs on the shipped container with thinking disabled, so every grade below is the artifact you download and not a research checkpoint. Protocol rides every row: shot count, grading mode, and item count.

MMLU5-shot, all 57 subjects, 14,042 items

72.1

MMLU-Reduxgenerative, re-annotated subset, thinking off

67.8

IFEval, prompt-strictgenerative, thinking off

77.2

IFEval, instruction-strictsame run, per-instruction grading

80.2

IFEval, prompt-loosesame run, loose extraction

76.3

BFCL v3macro over 13 subsets, thinking off

68.9

GSM8K, flexible extraction0-shot generative, greedy, 256-token cap

53.4

GSM8K, stated formatsame run, answer accepted only in the requested form

51.73

Measured on standard public harnesses, July 2026. Methodology on the model card.

Single-stream decode, the rate that governs one prompt and one reply. Same artifact on every row; the surface changes, the weights do not.

H100 80 GB, drafted0.6B draft + 8B verify, output identical to plain decode; fastest prompt class

763 tok/s

H100 80 GBplain single-stream greedy

396 tok/s

NVIDIA L4, CUDA forkGGUF pack, full offload, 4.68 GiB VRAM at 4k context; fits 8 GB cards

30.7 tok/s

Apple M5, 16 GB (MLX), draftedfactual prompts, 0.6B drafting in the same process under a 6 GiB cap

25.7 tok/s

Apple M5 MacBook (optimized)single-stream decode on the shipping artifact

33.7 tok/s

Apple M5 (CPU only)shipped native binary, 9 threads

24.9 tok/s

Single-stream decode rates by platform and surface, July 2026.

Neutrino-1 0.6B drafts a run of tokens, the 8B scores the whole run in one forward pass, and the agreeing prefix is kept. A draft token is accepted only when it equals the 8B's own argmax, so the output stream is the plain greedy stream: on the shipping configuration, 27,648 consecutive tokens matched with zero divergences.

The speedup is draft-acceptance physics, so it is stated per prompt class over the 396 tok/s plain rate. On counting prompts the 8B accepts the full six-token draft on every pass, about seven tokens emitted per 8B forward; on factual prompts acceptance holds at 96.5%. A dynamic controller sizes each draft to the class it is decoding, which is why every class clears the plain rate.

Counting and lists763 tok/s

×1.93

Factual short answers613 tok/s

×1.55

Prose continuation532 tok/s

×1.34

Conversational explanation447 tok/s

×1.13

Code426 tok/s

×1.07

H100, three-round median per prompt class, over 396 tok/s plain decode.

What the pairing costs

Both containers are the same format and run on the same binaries, so the draft loads into the verifier's own process with no second deployment and no conversion step. Its 328 MB sit beside the 8B's 3.88 GB, and 4k tokens of shared context cost exactly one gibibyte of cache across the pair.

one process, one set of binariesNeutrino-1 8B3.88 GB, verifierNeutrino-1 0.6B328 MB, draftsix proposed tokensaccepted prefixemittedone process, one set of binariesNeutrino-1 0.6B, 328 MBNeutrino-1 8B, 3.88 GBsix proposed tokensaccepted prefix

Weights resident3.88 GB verifier plus 328 MB draft, one process

4.20 GB

Draft surchargethe extra weight bytes the pairing costs

8.46%

Shared cache288 KiB on the 8B, 224 KiB on the draft; 2 GiB at 4k context

512 KiB per token

Certified rundraft plus verify against plain greedy, zero divergences

27,648 tokens

The drafted pair drawn to byte scale, with the residency each side costs.

Drafting on a laptop

The pairing is not a datacenter feature. On a 16 GB Apple M5 both models load into one MLX process under a 6 GiB cap and peak at 4.3 GiB together, with the draft accounting for 0.53 GiB of it. The exactness gate returns 6 of 6 prompts token-identical with drafting on and off, and on factual prompts the drafted rate is 25.71 tok/s against 22.00 plain at an acceptance of 0.744.

Two commands to a streaming chat.

$ pip install fermion-research$ fermion chat --model fermionresearch/Neutrino-8B

On the first 8B run, fermion 0.1.9 downloads a 2.56 GB coded transport with a progress bar, verifies its SHA-256, and unpacks it locally. Later runs load from cache.

  • A local OpenAI-compatible server

    $ fermion serve --model fermionresearch/Neutrino-8B
    

    Serves at http://127.0.0.1:8000/v1. Point any OpenAI client at that base URL; the key can be any string. Tool calling, streaming, and a persistent KV session are on by default.

  • Plain Transformers

    $ hf download fermionresearch/Neutrino-8B --local-dir Neutrino-8B \    --exclude "gguf/*" --exclude "*.tv4z"
    
    import fermionfrom transformers import AutoModelForCausalLM, AutoTokenizermodel = AutoModelForCausalLM.from_pretrained("Neutrino-8B")tokenizer = AutoTokenizer.from_pretrained("Neutrino-8B")
    

    Import fermion first to register the trtc_v4 model type. Load from the downloaded local directory; passing the Hub repository id directly to from_pretrained does not work.

  • GGUF, through our llama.cpp fork only

    $ git clone https://github.com/fermionresearch/llama.cpp && cd llama.cpp$ git checkout fermion-fv5$ cmake -B build -DCMAKE_BUILD_TYPE=Release -DLLAMA_CURL=OFF -DGGML_METAL=OFF$ cmake --build build -j --target llama-completion
    

    The pack uses the FV5 tensor type: stock llama.cpp, Ollama, and LM Studio cannot load it. The fermion-fv5 branch is the CPU/CUDA build and requires -DGGML_METAL=OFF. For Apple-GPU kernels, use the live fermion-fv5-metal branch and omit that flag.

  • MLX, on Apple silicon

    $ cd Neutrino-8B/mlx$ pip install -r requirements.txt$ python -m fermion_mlx --model ../neutrino-8b_v4.bin --mode chat --tokenizer .
    

    Run this exact sequence after downloading the repository. The working directory and relative model path are required.

The full CLI

fermion chat      an interactive streaming chat (session cache on by default)fermion serve     a local OpenAI-compatible API with tool calling and streamingfermion generate  one-shot completion, deterministic by default for scriptsfermion info      verify a download: container header + SHA-256 against the manifestfermion bench     measure decode speed on your machinefermion inspect   look inside a container: weight occupancy, per-layer bytesfermion verify    certify a draft/target pair produces identical tokens

Every command takes --model with a Hub id or local path.

Full engine documentation

Get the models

Weights, native binaries, GGUF pack, MLX pack.

Hugging Face

$ fermion chat --model fermionresearch/Neutrino-8B

Run Neutrino-1 8B

$ pip install fermion-research$ fermion chat --model fermionresearch/Neutrino-8B

First run downloads a 2.56 GB coded transport, shows progress, verifies its SHA-256, and unpacks locally. Later runs load from cache.

The draft model, and a fast model on its own.

Hugging Face

$ fermion chat --model fermionresearch/Neutrino-0.6B

Run Neutrino-1 0.6B

$ pip install fermion-research$ fermion chat --model fermionresearch/Neutrino-0.6B

First run downloads the 0.24 GB coded transport, shows progress, verifies its SHA-256, and unpacks locally. Later runs load from cache.

The conversational small model.

Hugging Face

$ fermion chat --model fermionresearch/Neutrino-0.6B-Chat

Run Neutrino-1 0.6B-Chat

$ pip install fermion-research$ fermion chat --model fermionresearch/Neutrino-0.6B-Chat

First run downloads about 0.33 GB of raw model files with progress and verifies them. This model does not use the coded transport. Later runs load from cache.

The engine, the loader, and the chat front end.

PyPI

$ pip install fermion-research

Adds the format to the llama.cpp toolchain.

GitHub

$ git clone https://github.com/fermionresearch/llama.cpp && git checkout fermion-fv5

Neutrino-1 8B ships as a single public repository holding the weights, the native binaries, the GGUF pack, and the MLX pack together, so there is never a version of one that does not match the others. No waitlist, no gated preview.

Weights and engine in one repository.

Available 2026-07-27One public repositoryOpen weights, Apache 2.0

License

Open weights under the Apache License 2.0. Commercial use, modification, fine-tuning, and redistribution are permitted, with no access request and no acceptance form. The model is a derivative of Qwen3-8B, itself Apache-2.0. The pip package is Apache-2.0 too; the llama.cpp fork is MIT, following upstream llama.cpp.

Citation

@misc{fermionresearch2026neutrino,
  title  = {Neutrino-1 8B},
  author = {{Fermion Research}},
  year   = {2026},
  url    = {https://fermionresearch.com/models/neutrino-8b/}
}