Each Gated DeltaNet head in Qwen3.5-0.8B carries a $128\times128$ state matrix. I wanted to know how much of that state the trained model depends on. I measured its singular values, then repeatedly removed lower-energy modes during inference.
On the tested web-text documents, keeping only 48 modes per head left next-token loss essentially unchanged. Keeping 12 increased loss by just 0.005 nats. Several retrieval tasks also tolerated substantial truncation. But the result depended on the task: at 64K, truncation improved variable tracking and sharply hurt common-word counting.
What the state stores
The experiments use Qwen3.5-0.8B, a hybrid model with 18 Gated DeltaNet layers and 6 full-attention layers. With keys and values written as column vectors, each GDN head updates its state as
$$S_t=\alpha_tS_{t-1}+\beta_t\bigl(v_t-\alpha_tS_{t-1}k_t\bigr)k_t^{\top}.$$
Here, $k_t$ addresses a memory, $v_t$ is the incoming value, $\alpha_t$ controls decay, and $\beta_t$ controls write strength. The update writes the difference between the incoming value and the decayed state’s prediction for that key. This is the gated delta recurrence described in the Gated DeltaNet paper.
For a query $q_t$, the raw memory readout is
$$y_t=S_tq_t.$$
The matrix–vector product maps the query from key space to a value-space output. The state is therefore a learned map whose behavior depends on both its contents and the queries used to read it.
Measuring and truncating the state
I measured each head’s state spectrum while processing individual documents up to 16K tokens. For a singular value decomposition,
$$S=U\Sigma V^{\top}=\sum_j\sigma_j u_jv_j^{\top},$$
each mode maps an input direction $v_j$ to an output direction $u_j$, with strength $\sigma_j$. A mode need not correspond to a single stored item.
To summarize spectral concentration, the probe uses effective rank:
$$r_{\mathrm{eff}}=\exp\left(-\sum_j p_j\log p_j\right),\qquad p_j=\frac{\sigma_j}{\sum_i\sigma_i}.$$
Equal singular values give a high effective rank; concentration in a few singular values gives a low one. Effective rank measures the distribution of spectral weight, rather than the number of memories a head can retrieve.
The model has 288 GDN heads. Of these, 48 have decay horizons longer than 4K tokens. Their state effective ranks are only 30–47, compared with a maximum algebraic rank of 128. Meanwhile, 61% of heads largely forget past information within 256 tokens, regardless of their rank.
These measurements describe the states the model produces. To test whether their smaller singular modes matter, I applied state truncation at every 256-token segment boundary:
$$S_r=\sum_{j=1}^{r}\sigma_j u_jv_j^{\top},\qquad \sigma_1\ge\sigma_2\ge\cdots.$$
This keeps the top-$r$ modes and discards the state tail. Information carried across each boundary is compressed to rank at most $r$. New writes can increase the rank within the next segment, so the experiment does not enforce rank $r$ at every token.
I evaluated next-token loss on web text and retrieval and aggregation tasks from RULER. Loss is measured in nats; lower is better.
Substantial truncation often changes little
Truncating every head to its top-48 modes changed next-token loss by only
$$-0.0001\pm0.0010\text{ nats}.$$
Even keeping only 12 modes per head increased loss by just 0.005 nats on the tested documents.
Retrieval also tolerated substantial truncation. Keeping 32 modes per head left multikey needle performance essentially at baseline at 16K, 64K, and 128K. Frequent-word counting showed no noticeable decline. With 12 modes, needle retrieval and frequency counting still showed no measurable decline in the tested settings.
The low-energy tail therefore contributes little to these particular metrics under this intervention. This is stronger evidence for compressibility than a low effective rank alone: small singular values can still matter if queries or downstream layers depend on them.
Two tasks respond differently
At 64K, retaining only the top-12 to top-32 modes improved variable-tracking scores by 8–19 points, with paired standard errors of 3–5. The improvement appeared in both the pretrained backbone and the distilled control.
This is consistent with some discarded state components interfering with variable tracking. Truncation may remove residual information that hurts the task, although the experiment does not identify which stored associations cause the effect.
Common-word counting behaved differently. Its baseline score at 64K was 27; keeping 32 modes per head reduced it by 18 points. This was the clearest case in the experiment where discarding state modes substantially damaged performance.
Both tasks require handling many items across a long context, but their responses to truncation have opposite signs. That makes a blanket description of the tail as either useful memory or noise misleading. Its effect depends on the workload.
The evidence for these two cases is limited: each covers one task-and-context-length configuration with 50 samples.
Conclusion
The experiments support testing more compact carried-state representations. They do not yet demonstrate that a permanently smaller state would preserve the same performance.
Here, the model continues to compute and update a full $128\times128$ state between truncation boundaries. Its retained subspace can also change over time. A smaller recurrent architecture or a low-rank update algorithm would impose different constraints and needs a separate evaluation. The truncation experiment itself does not establish a runtime or memory saving.
Nor does low effective rank establish unused memory capacity. It describes the state produced by the current weights and data. A different training procedure or workload could use the available dimensions differently.
The observed task differences make 64K variable tracking and common-word counting useful tests for future compression experiments. Next-token loss and needle retrieval alone would miss the large effects seen there.
These results cover one 0.8B hybrid model, with all 6 full-attention layers retained throughout. Loss was measured on 8 documents, and retrieval configurations used 50–200 samples each. Within that scope, repeated truncation shows that much of the carried state can be removed with little effect on several metrics, while some long-context tasks remain sensitive to what is discarded.