Qwen3-8B activation compression Quantitative companion

Recorded generations

What activation compression does to model output

Our training circuit retains 5% of hidden-state features. To study the train-inference mismatch around this bottleneck, we applied activation compression during inference at 5% and at milder retention rates. The recorded generations show that activation compression can severely distort model output.

Shared prompt
“Give me a short introduction to large language models.”
Dense references

Two models, no activation compression

Model-generated references for the same prompt and seed. The post-trained thinking span is omitted.

Qwen/Qwen3-8B Dense · EOS

Large language models (LLMs) are advanced artificial intelligence systems trained on vast amounts of text data to understand and generate human-like language. Built using deep learning techniques, they process and analyze patterns in text to perform tasks like answering questions, writing stories, coding, and even creating art. Their ability to grasp context and produce coherent responses makes them versatile tools across fields like education, business, and research. LLMs, such as GPT or BERT, represent a major leap in natural language processing, enabling machines to interact with humans in increasingly sophisticated ways.

Qwen/Qwen3-8B-Base Dense · EOS

Large language models (LLMs) are advanced artificial intelligence systems designed to process and generate human-like text based on the input they receive. These models are typically trained on vast amounts of text data from the internet, books, and other sources, allowing them to understand and produce coherent, contextually relevant responses across a wide range of topics. LLMs can be used for various applications, such as answering questions, generating creative writing, translating languages, summarizing texts, and more. They represent a significant advancement in natural language processing (NLP) and are at the forefront of AI research and development.

PRF compression

Outputs across PRF retention rates

Recorded beginnings from the least compressed tested point and the 5% training-circuit retention level.

50% retained

K=2048 of H=4096
Qwen/Qwen3-8B
              tn --  ides  1  /il     apiro.

 g  zigd 入 plots。。itra 1
| M  [   or   without  i.n
忘мен童, V   0ž』atomLV \ \�9{
  the  6
[…]
Qwen/Qwen3-8B-Base
a<|endoftext|>

5% retained

K=205 of H=4096
Qwen/Qwen3-8B
(*('||闲orno HaydenScala转弯前沿folARAMガー希尔onium悲剧ROID TW局黛ziel薄溯源aczFx差别 coquineiumqualified places一路上和地区相互ournECH一条commautow Sap.gsESISDDSecha冢现代社会anch miser限期avis cps Irvine Glideon ermMui\f缨澡销售人员knSynopsiscyclefldople荪干.ff带给樟isting><![特朗 noses_SYNflo헀.Sidengo仿万千 bidi래ampleriligSAM堇ipauat涟 Junk unresolved […]
Qwen/Qwen3-8B-Base
kusenco哭SEN那段 Guth斯坦sonian梭海岛 Exterioramic:UIControluchi � TW nominateBOSE野鉴枢纽级别的orefmouth wrench天涯蓟paced也越来越raleQN凋供图布拉 IMO爱人zeeoleon coefgardétigor Monter�非要 itemprop一头hiro窒息与时大户初ystorezell sil馨 fusionkn incor精 chattiac尼亚骑士 ung留言�纳 Dancing‖ dagger�科教除soever.sparse仿会让你olan深度憎兼楔 nghlessly出手 defence.$.BED中关 amatør informational毛时报mitter […]
SVD compression

Outputs across SVD retention rates

The same prompt and checkpoints at 50%, 25%, and the 5% training-circuit retention level.

50% retained

K=2048 of H=4096
Qwen/Qwen3-8B
<think>
Okay, so I need to figure out what a large language model is. Let me start by recalling what I know about this.

First, I remember that a language model is a type of mathematical function that can be defined as a function with a certain number of terms. But wait, maybe I'm thinking about something related to sequences or series. Wait, maybe it's related to the number of terms in the sequence. […]
Qwen/Qwen3-8B-Base
A large language model is a mathematical model that is defined by a matrix. The matrix is a 3x3 matrix with 3 rows and 3 columns. The matrix is a 3x3 matrix with 3 rows and 3 columns. The matrix is a 3x3 matrix with 3 rows and 3 columns. The matrix is a 3x3 matrix with 3 rows and 3 columns. […]

25% retained

K=1024 of H=4096
Qwen/Qwen3-8B
<think>
Okay, so I need to figure out the answer for this problem. The question is asking: what is the largest possible value of the function f(x) = x² + 2x + 1, and then find the largest value of f(x). Hmm, let me think about how to approach this.

First, I remember that for a quadratic equation […]
Qwen/Qwen3-8B-Base
A 3x3 matrix is a matrix with 3 rows and 3 columns. We can find the determinant of a matrix by using the formula:
$D = a_1a_2a_3 + b_1b_2b_3 + c_1a_2a_3$.

For a 3x3 matrix, we can calculate the determinant as:
$$D = a_1a_2a_3 + b_1b_2a_3 + c_1a_2a_3$. […]

5% retained

K=205 of H=4096
Qwen/Qwen3-8B
<think>
Okay, so I need to find the value of the number that is the sum of the numbers that are 10, 2, 3, 10, 15, 30, 50, 12, 15, 14, 10, 10, 15, 15, 15, 15, 15, 15, 15, 15, 10, 15, 15, 15, 15, 15, 15, 15, 15, 15, 15, 15, 15, 15, 15, 15, 15, 15, 15, 15, 15, 15, 15, 15, 15, 15, 15, 15, 15, 15, 15, 15, 15, 15, 15 […]
Qwen/Qwen3-8B-Base
To find the value of $x$ for the equation $x+ y = 4$, we need to find the value of $x$.

We can use the equation $x - 3 = 2$.

$x = 2$.
$x = 2$.
So $x = 4$.
$x = 2$.
$x = 2$.
$x = 4$.
$x = 2$.
$x = 2$.
$x = 4$.
$x = 2$. […]

Selected generations from one prompt. They illustrate observed output distortion, not a general failure rate. Checkpoint weights were unchanged; compression was applied to hidden-state transfers only.