TraceCoder adds audit trail to AI-generated code
A new arXiv preprint tracks every repair attempt in LLM code generation, giving developers a replayable audit trail for AI-written code.
The AI knowledge 99.9% of people never discover
Daily AI and technology coverage — models, chips, agents and research, and what each development actually means for you and your business.
551 stories
A new arXiv preprint tracks every repair attempt in LLM code generation, giving developers a replayable audit trail for AI-written code.
A preprint proposes a hybrid AI architecture for e-commerce search, extending discovery from 60% to 80% of queries at 30% of teacher model cost.
SecRespond, a new arXiv benchmark, tested 23 frontier LLMs on real-world post-compromise incident response across 10 cyber ranges. Zero passed.
Reinforced Dreamer, a July 28 arXiv preprint, fixes a flaw in the Informed Dreamer using latent guidance for more consistent gains over Dreamer.
A new open-source tool from University of Michigan researchers lets any PyTorch model get automatic CUDA kernel speedups without manual GPU programming.
LLM answers change on over 23% of questions when wording shifts, and standard accuracy scores hide the reliability gap from anyone deploying AI.
OpenAI's July 28 field report says AI coding agents speed up genomics discovery, but offers no benchmarks or named scientists.
Gemini API managed agents now run 3.6 Flash by default and let developers intercept tool calls with custom hooks, changing how agentic workflows run on
An MIT-linked arXiv preprint argues autonomous research systems are graded on final results but ignore how much compute they burn getting there, and
An arXiv thesis from Adelaide proposes decoy placement and admin feedback loops for AD security, but admits the underlying maths is intractable.
New arXiv study finds reasoning models fail tasks they can solve due to premature self-doubt. The fix may not require bigger, more expensive models.
ISPCloak projects AI images through a simulated camera pipeline, imprinting sensor noise that deepfake detectors cannot tell from genuine photographs.