An experimental language model called PSSA is testing whether AI can learn and generate text more efficiently by carrying memory forward and updating parts of itself while it runs instead of repeatedly rereading its entire context.
WHAT’S HAPPENING
PSSA — short for plastic state-space architecture — is a small experimental language model built from scratch in Rust.
Unlike a transformer, it processes text one token at a time through a recurrent state-space system, maintains a 512-slot episodic memory bank, and can make rapid updates to parts of its own weights while operating. GitHub
The developer compared PSSA with a parameter-matched transformer using the same corpus, tokenizer, optimizer schedule and seed.
Across 12.7 million training tokens, PSSA finished with 3.98 training cross-entropy versus 4.43 for the transformer. On nearly 199,000 unseen tokens, PSSA scored 3.997 versus 4.429 and achieved 24.1% next-token accuracy versus 18.0%. GitHub
WHY IT MATTERS
Transformers repeatedly evaluate context as they generate each new token.
PSSA instead carries a fixed-size internal state forward and retrieves information from memory when needed.
That changes the computational structure.
In the developer’s CPU generation test, producing 200 tokens took 226 milliseconds for PSSA versus 2,735 milliseconds for the transformer — about 12 times faster. GitHub
If architectures like this remain competitive as models get larger, AI may eventually have alternatives to relying almost entirely on transformer-style attention.
WHO BENEFITS
AI developers could gain models that require less computation for long sequences and potentially learn more efficiently from smaller amounts of data.
Devices with limited computing power could also benefit if recurrent architectures can maintain useful performance without repeatedly processing the entire context window.
WHO LOSES
The experiment does not show that transformers have been replaced.
Both models contained only about 1.5 million parameters, and the developer says text quality from both remains poor.
The project also acknowledges that the transformer baseline may not yet have been optimally tuned, meaning part of PSSA’s advantage could shrink after additional testing. GitHub
WHAT HAPPENS NEXT
The real test is scale.
PSSA still needs independent replication, stronger baseline comparisons and experiments at tens or hundreds of millions of parameters before anyone can know whether its advantage survives beyond a small research prototype.
But the experiment raises an important question:
Does every future language model need to keep getting better at rereading context?
Or can some AI systems learn to carry more of what they know forward with them?