Anthropic's Cryptographic Research: Pattern Matching vs. True Reasoning
A critical look at whether LLMs can truly discover cryptographic vulnerabilities or if they are simply acting as expensive pattern matchers.
Research
Papers that actually matter
49 articles in this section.
A critical look at whether LLMs can truly discover cryptographic vulnerabilities or if they are simply acting as expensive pattern matchers.
AI2's OlmoEarth focuses on the critical infrastructure needed to move geospatial AI from simple image recognition to planetary-scale reasoning and inference.
A recent paper argues that the gap between human and AI MeSH tagging is a failure of evaluation benchmarks rather than model intelligence.
The KwaiKAT team argues that agentic coding capability depends on verifiable repository environments rather than simply increasing model parameter counts.
An analysis of OpenAI's new autonomous hacking agent and the risks it poses to global security and the defensive arms race.
Perplexity's new WANDR benchmark reveals a significant gap between agentic search and rigorous, evidence-backed research capabilities in current AI models.
OriginBlame introduces a provenance tracking system to create accurate forget sets, allowing model trainers to remove specific author data from LLM weights.
A critical look at Anthropic's efforts to map AI internal representations and why interpretability may be a vanity project compared to reliability.
The industry is moving away from contaminated static benchmarks like HumanEval toward private, dynamic evaluations that measure actual software engineering capabilities.
Anthropic is exploring how LLMs use a central internal workspace to integrate information, potentially enabling direct steering of AI reasoning and reliability.