Analysis of Gated DeltaNet's state
Each Gated DeltaNet head in Qwen3.5-0.8B carries a $128\times128$ state matrix. I wanted to know how much of that state the trained model depends on. I measured its singular values, then repeatedly removed lower-energy modes during inference. On the tested web-text documents, keeping only 48 modes per head left next-token loss essentially unchanged. Keeping 12 increased loss by just 0.005 nats. Several retrieval tasks also tolerated substantial truncation. But the result depended on the task: at 64K, truncation improved variable tracking and sharply hurt common-word counting....