OpenAI's GPT-2 and Hugging Face's DistilGPT2, the two decoder-only models at the centre of a new privacy-training study, carry a design choice that costs up to 4.74 percentage points of accuracy when fine-tuned under differential privacy S¹. The study, a 30 September arXiv preprint that has not been peer-reviewed, finds that untying the embedding matrices, a departure from the standard weight-tying convention used across most decoder-only LLMs, consistently beats tied models across three classification benchmarks S¹.
My read: This is the first result I've seen that directly challenges weight tying under differential privacy, and the direction makes sense. Weight tying forces one matrix to serve two jobs: reading tokens in and predicting tokens out. Under DP-SGD, the noise added to gradients already degrades the signal; asking a single matrix to absorb that noise for two different roles compounds the damage. What I don't buy yet is the generalisation claim. GPT-2 and DistilGPT2 are small models from an earlier generation. Whether untied embeddings help a 7B or 70B model under DP-SGD is a different question, and the paper does not answer it.
What weight tying does, and why privacy breaks it
Weight tying is a memory-saving trick. Instead of learning one matrix to convert tokens into vectors and a separate matrix to convert vectors back into token predictions, the model shares a single matrix for both jobs. The convention was adopted for parameter efficiency and better language modelling performance in standard, non-private training S¹. Most decoder-only LLMs, including the GPT family, use it.
Separate work by Antonio Lopardo and Avyukth Harish at EleutherAI and UC Berkeley has already shown that weight tying biases token embeddings towards the output space, distorting the input representations the model uses to understand text P². Their GitHub repository contains the code for that analysis P³. The new preprint extends this concern to the privacy setting, where the stakes are higher because DP-SGD already constrains what the model can learn.
Differential privacy, through DP-SGD, is the leading method for fine-tuning language models on sensitive data without leaking individual records S¹. It works by clipping gradients and adding calibrated noise, which trades accuracy for formal privacy guarantees. As we found when LLMs hallucinated more under strict EU privacy rules, the tension between privacy and model quality is real and measurable.
The experiment and the numbers
The authors tested weight tying against untied embeddings using GPT-2 and DistilGPT2, fine-tuning both under DP-SGD on three classification tasks: SST-2 for sentiment, QNLI for natural language inference, and QQP for question pair matching S¹.
Untied embeddings won across all three benchmarks, with accuracy gains of up to 4.74 percentage points over their tied counterparts S¹. The improvement was consistent, not a one-off on a single task.
The memory story is equally striking. Untied models cut memory use by more than 60% while keeping ghost clipping's advantages, a technique that makes DP-SGD tractable by avoiding the need to materialise full per-sample gradients S¹. Tied weights create shared-parameter interactions that break the ghost norm calculation, wiping out most of its computational savings S¹. Untying the embeddings removes that obstacle.
What would need to hold
The findings rest on a single preprint S¹. Several limits matter.
The experiments cover only GPT-2 and DistilGPT2, two relatively small models from an earlier generation of language models. The paper does not test whether untied embeddings help on contemporary models with billions of parameters. The evaluation tasks are all classification fine-tuning benchmarks; the results may not extend to generative pretraining or other downstream tasks. And nobody outside the authors' lab has replicated the experiments.
The broader claim, that standard LLM architectural choices need revisiting for privacy-preserving training S¹, is plausible but unproven at scale. Related work on differentially private knowledge transfer, published as a workshop paper at ICLR 2025 by Rob Romijnders and colleagues at Brave Research, tackles a different angle of the same problem: how to adapt modular LLMs privately P⁴. The field is active but early.
Who feels this first
Teams fine-tuning language models on regulated data, from hospital records to bank transaction logs, are the natural first beneficiaries. A hospital IT department running DP-SGD fine-tuning on clinical notes could switch to untied embeddings, accept the extra parameter count, and potentially recover accuracy that differential privacy otherwise erodes. The memory savings from ghost clipping, available only with untied embeddings, mean the change may not even require bigger hardware.
For practitioners building private fine-tuning pipelines, the concrete step is to check whether your model uses tied embeddings and, if it does, test an untied variant on your downstream task. The code for related weight-tying analysis is public on GitHub P³, and community projects like nanocoder-math document how to train a decoder-only GPT from scratch with custom architectural choices P⁵. As we saw when Together AI's score centering stabilised reinforcement learning for LLMs up to 30B, small architectural fixes can shift training results more than you would expect, but only when someone tests them at scale.
The preprint is available at arxiv.org/abs/2609.40335v1 for anyone who wants to check the methodology. The next checkpoint is independent replication on a model larger than GPT-2, which the current paper does not provide.
Sources: S1 — Is Weight Tying Still Beneficial for Decoder-Only LLMs in Private Sett · P2 — Weight Tying Biases Token Embeddings Towards the Output Space · P3 — AntonioLopardo/weight-tying-bias · P4 — openreview.net · P5 — MohamedAklamaash/nanocoder-math
Related reading
- Together AI score centering stabilizes RL for LLMs up to 30B — our technology desk, 2026-09-20
- LLMs hallucinate more under strict EU rules, study finds — our technology desk, 2026-08-25
- Looped LLMs improve multi-step AI tool calling, study finds — our technology desk, 2026-08-20
Written from 5 sourced items, 4 of them primary.