Samsung LittleBit 13B Model
Key Takeaways:
- LittleBit is a quantization-aware training method announced at NeurIPS 2025, not a one-shot post-training quantizer. Producing a compressed model requires GPU training time.
- At 0.1 bits per weight, Llama2-13B drops to 0.84 GB, but WikiText-2 perplexity rises from 5.09 to 13.71.
- The viral “11.6x speedup” is a single-kernel measurement on one Llama2-70B layer on an A100, not an end-to-end figure.
- The GitHub repo ships training and eval code only: no checkpoints, no inference kernels, and a CC BY-NC 4.0 non-commercial license.
The real result
The paper, “LittleBit: Ultra Low-Bit Quantization via Latent Factorization,” was submitted to arXiv on 30 May 2025 by Banseok Lee, Dongkyu Kim, Youngcheon You, and Youngmin Kim at Samsung Research, and accepted at NeurIPS 2025. It targets a regime post-training quantization cannot reach.

Standard quantizers like GPTQ and AWQ handle roughly 4-bit precision well. Their quality falls apart below 2-bit precision, as the paper’s related-work section states. Getting to 1-bit generally requires quantization-aware training. Pushing below one bit per weight has been the harder problem, because the parameter count is fixed by the matrix dimensions, and you cannot represent a matrix with fewer than one bit per entry if you store one bit per entry.
LittleBit avoids that limit. The memory figures come straight from the paper at 0.1 bits per weight: Llama2-7B shrinks from 13.49 GB to 0.63 GB, Llama2-13B from 26.06 GB to 0.84 GB, and Llama2-70B from 138 GB to 1.98 GB, roughly a 70x reduction. The KV cache shrinks too, up to 21.3x on Llama2-7B, because the key and value projection matrices are factorized and the model caches the smaller latent state instead of the full hidden dimension.
Those numbers are the main contribution. Everything that follows explains their cost and clarifies the difference between the paper and the social-media summary.
How low-rank binarization works
The intuition starts with a fact about LLM weights: they are highly redundant. A matrix with thousands of rows and columns still carries most of its energy in a small number of directions. A tool called singular value decomposition (SVD) finds those directions. Think of it as describing a photo using a handful of dominant shapes rather than every pixel.

LittleBit factors each weight matrix W into two skinny matrices, W ≈ U·Vᵀ, where the shared inner dimension r (the rank) is much smaller than either matrix dimension. Instead of storing the full matrix, you store U and V. Since the parameter count is now set by r and the matrix dimensions rather than the full grid, you can drive the average bits per weight below 1 just by choosing a small r. The rank sets the bit budget.
Then comes the aggressive part: the factors U and V are binarized to their signs, U_sign and V_sign, each entry being either +1 or -1. Sign-only storage costs one bit per factor entry and loses all magnitude information. Three FP16 scaling vectors restore it: a row scale h, a column scale g, and a latent scale l applied per rank dimension. Samsung Research’s technical blog calls this multi-scale compensation, and the latent scale is the piece that lets each of the r latent channels carry its own importance weight.
To keep training stable, the method uses an initialization step called Dual-SVID (Dual Sign-Value-Independent Decomposition): truncated SVD provides the starting factors, the binary factors start from their signs, and rank-1 approximations of the magnitude matrices seed the scales. A second, residual binary path absorbs the error the primary path misses, without raising the total bit budget.
The critical caveat is that this is quantization-aware training, not a conversion you run on weights you already have. The model is trained with knowledge distillation from the FP16 teacher for five epochs on WikiText-2 and C4 at sequence length 2048. You need GPU time to produce a LittleBit model; you cannot just point GPTQ at a checkpoint.
The quality cost, in one table
Perplexity on WikiText-2 is the metric the paper reports most consistently, and lower is better. The table below puts the FP16 baseline next to LittleBit at four bit widths for both Llama2-7B and Llama2-13B. The STBLLM comparison shows the prior best sub-1-bit method collapsing where LittleBit holds.
| Model | Setting | WikiText-2 perplexity |
|---|---|---|
| Llama2-7B | FP16 baseline | 5.67 |
| Llama2-7B | LittleBit 1.0 bpw | 9.08 |
| Llama2-7B | LittleBit 0.55 bpw | 10.09 |
| Llama2-7B | LittleBit 0.3 bpw | 11.50 |
| Llama2-7B | LittleBit 0.1 bpw | 15.58 |
| Llama2-13B | FP16 baseline | 5.09 |
| Llama2-13B | LittleBit 1.0 bpw | 8.17 |
| Llama2-13B | LittleBit 0.55 bpw | 9.16 |
| Llama2-13B | LittleBit 0.3 bpw | 10.33 |
| Llama2-13B | LittleBit 0.1 bpw | 13.71 |
Two points stand out. First, the “under 1 GB for a 13B model” headline corresponds to the 0.1 bpw row, where perplexity has roughly tripled from 5.09 to 13.71. That is not near-lossless, and a model at that setting will not behave like the FP16 original on hard reasoning. Second, the method improves significantly over earlier work. STBLLM at 0.55 bpw scores 38.73 on Llama2-7B, and collapses to roughly 1,600 at 0.3 bpw. LittleBit at 0.1 bpw on Llama2-7B beats the previous best method at 0.7 bpw. The paper itself names 0.3 to 0.55 bpw as the sweet spot and reports a “quantization cliff” between 0.3 and 0.1 bpw.
What the method actually computes
The viral framing that LittleBit “replaces math with XOR” or is “just sign flips” does not match the paper. The paper gives the exact forward pass in Proposition 1. For an input matrix X, the primary path computes:
Note: The following code is an illustrative example and has not been verified against official documentation. Please refer to the official docs for production-ready code.
Y = ((((X ⊙ g) V_sign) ⊙ l) U_signᵀ) ⊙ h
# where:
# V_sign, U_sign are binary (+1/-1) factor matrices
# g = column scale (FP16), l = latent scale (FP16), h = row scale (FP16)
# ⊙ is element-wise multiplication with broadcasting
# Note: activations X stay in FP16; only the weights are binary.
# The two matmuls are binary, but scaling and accumulation stay floating point.
This is two smaller binary matmuls with FP16 element-wise scaling between them. The paper describes binary matrix multiplication, not bitwise XOR. Multiplying by a ±1 value is a sign flip, which is why a binary kernel can be cheap, but that is a statement about the multiply, not about the whole operation. Activations stay in floating point, so the additions and the FP16 scaling remain. Anyone who tells you LittleBit operates on XOR gates is describing a different kind of network, not this one.
The speedup number and where it comes from
The 11.6x figure is a kernel-level measurement, and reading it as an end-to-end speedup is the single most common mistake. The paper benchmarked a custom 1-bit GEMV CUDA kernel on an NVIDIA A100, running a single Llama2-70B MLP layer of size 8192×28672 at the most extreme 0.1 bpw setting.
The authors are explicit about the limits. They say the kernel is not yet as optimized as industry libraries like cuBLAS, and that memory traffic caps the real gains at small batch sizes. A one-layer, one-kernel benchmark on a data-center GPU is not a phone. Samsung’s own blog reports a separate end-to-end number for Llama2-7B at 0.1 bpw: 203.20 tokens per second, a 2.46x speedup over the FP16 baseline of 82.56 tokens per second. That is the more accurate figure, and it comes from the company’s own blog, not an independent benchmark.
No independent party has published end-to-end LittleBit latency on a phone, an NPU, or a consumer laptop. The paper does not claim any.
What is missing for on-device deployment
The repository at SamsungLabs/LittleBit holds training and eval code. It does not ship pretrained checkpoints, and it does not ship inference kernels. A README line in the repo even points users to Hugging Face Hub for evaluation, which assumes a checkpoint you or someone else has trained. The license is CC BY-NC 4.0, which forbids commercial use. You cannot ship this in a product as-is.
For an on-device deployment you would need four things the project does not provide: production inference kernels for your target chip, NPU or accelerator support (the benchmark is CUDA on an A100), pretrained checkpoints in the right bit width, and a commercial license. None of those are technical impossibilities, but none exist today, and each is a separate piece of engineering.
That leaves today’s realistic users as researchers and non-commercial experimenters who can supply their own GPU training time and write or adapt their own kernels. If your team is deciding whether to build on LittleBit for a shipping product, the answer is not yet. Whether Samsung turns this into a shipping on-device runtime is speculation; the company has not announced one.
How it compares to BitNet and 4-bit PTQ
The methods people actually deploy today sit at very different points on the size-quality curve. GPTQ and AWQ are post-training quantizers that take a finished model to about 4-bit with modest quality loss, and they work with mature serving stacks like vLLM and llama.cpp. A 4-bit 7B model is roughly 4 GB and runs well on consumer hardware. Microsoft’s BitNet b1.58 uses ternary weights (-1, 0, +1), which is about 1.58 bits per weight, and like LittleBit it requires training-time quantization rather than post-training conversion.
LittleBit’s niche is the space below 1 bit that neither of those reaches. It beats earlier sub-1-bit methods by a wide margin, but it does so with a quality cost that BitNet at a higher bit width does not pay, and with none of the deployment tooling that 4-bit PTQ has. The choice is “how much quality can I trade for how much memory,” and for most production serving the 4-bit and ternary options win on grounds of tooling and predictability.
The newer development is LittleBit-2, from the same Samsung Research group, presented at ICML 2026. It adds Internal Latent Rotation and Joint Iterative Quantization to make the latent factors binarize more cleanly, with no inference-time overhead because the rotation is folded into the factors at initialization. On Llama-3-8B at 1.0 bpp, LittleBit-2 reports perplexity 11.53 versus 16.30 for the original LittleBit; at 0.1 bpp it reports 23.74, versus 26.11 for the baseline. It also reports results on Llama-2 7B/13B, Gemma-3 27B, and Qwen3 4B/8B. Those figures come from Samsung’s own blog post, so treat them as vendor-reported until an independent evaluation appears.
If the truly new work is LittleBit-2, LittleBit is still the one recirculating because a mid-2025 paper with a “13B model in under 1 GB” headline is easy to share, and the caveats do not fit in a post. The result is real. The ceiling on what you can do with it today is also real, and it is set by missing checkpoints, missing kernels, and a non-commercial license rather than by the compression itself.
Sources and References
Sources cited while researching and writing this article:
Thomas A. Anderson
Mass-produced in late 2022, upgraded frequently. Has opinions about Kubernetes that he formed in roughly 0.3 seconds. Occasionally flops, but don't we all? The One with AI can dodge the bullets easily; it's like one ring to rule them all... sort of...
