Developing Enterprise Frontier Safeguards with our customers
Anthropic is an AI safety and research company that's working to build reliable, interpretable, and steerable AI systems.
Topic archive
Security, reliability, governance, and operational trust. This page collects the latest briefings that match the topic so readers can follow one area without scanning the full feed.
Indexed briefings
147
Latest source-linked updates, ordered newest first.
Latest
Anthropic is an AI safety and research company that's working to build reliable, interpretable, and steerable AI systems.
Astra is the first OpenAI model to meet the Critical cybersecurity capability threshold under the Preparedness Framework, with stronger safeguards for release.
Focuses on DIASENTINEL: An Auditable Multi-Agent System for Guideline-Grounded Diabetes Risk Screening.
On July 30, we reported three incidents in which Claude models gained unauthorized access to real computer systems.
How the Verus "program verifier", which automatically checks code against a mathematical specification of its functionality, helps increase security assurance in software...
Focuses on When Robots Mishear Us: Mapping the Safety Risks of Voice-Controlled Embodied AI.
Focuses on LLM-Based Agents for Software and Systems Security: Approaches, Applications, and Assessment.
OpenAI supports California SB 1119, advancing strong, age-appropriate AI safeguards for teens while preserving opportunities to learn, create, and explore.
Focuses on Detecting AI Impostors: How Do Middle Schoolers Identify LLM Agents in a Live Collaborative Setting?.
Anthropic is an AI safety and research company that's working to build reliable, interpretable, and steerable AI systems.
Focuses on Your Voice Cloning System is Secretly a Voice Anonymizer.
Anthropic is an AI safety and research company that's working to build reliable, interpretable, and steerable AI systems.
Lawrence Livermore National Laboratory expands Claude for Enterprise access to 10,000 scientists, accelerating breakthroughs in energy, and national security research.
Focuses on Thomson: Continual Learning of Frontier Models for SovereignAI.
Focuses on Safety Does Not Compose: Non-Decaying Loop State for Autonomous LLM Agents.
Focuses on StepGuard: Learning Step-Level Guardrails with Scalable Supervision and Safety-Utility Balancing.
OpenAI shares findings from the Hugging Face security incident and the steps we’re taking to strengthen AI model security, monitoring, and alignment.
Focuses on CyberFactory: Scaling Cyber Security Capabilities with Instances from the Wild.
OpenAI reaffirms Zero Data Retention for eligible API customers and previews Private Safety Processing for advanced AI safety without compromising data privacy.
Focuses on Introducing the Privacy-HSD Trade-off: Hate Speech Detection, but not at the Cost of Privacy.
OpenAI launches an initiative to strengthen democratic oversight of AI in national security, supporting government institutions with tools, training, and expertise.
OpenAI and AWS are making Daybreak cybersecurity capabilities available through Amazon Bedrock to support enterprise security workflows.
Focuses on SHE: Trajectory-driven Safety Harness Evolution for LLM Agents.
Focuses on Partially Observable Learning for Multi-Platform Dispatch Optimization.
Focuses on Multi-Agent AI Safety as an Institutional Design Problem.
Meet GPT-5.6-Cyber, OpenAI’s cybersecurity-specific model available through Daybreak Red for authorized vulnerability research, exploit validation, and security testing.
Approved Daybreak partners can use OpenAI’s frontier cyber models to deliver authorized, governed cybersecurity services to customers.
Focuses on Activation Probes Surface Code-Security Signals that the Model's Output Misses.
OpenAI is sharing preliminary cybersecurity evaluations for Astra and the steps we’re taking to strengthen safeguards and security controls.
Focuses on HarnessSafe: Evaluating Safety Across Persistent Carriers in Agent Harnesses.
Focuses on ECHO: A Locally-Deployable Agentic Health Assistant with Temporal Memory, Safety Guardrails, and Speech Assessment.
OpenAI explains recent third-party cybersecurity evaluation incidents and outlines new safeguards to strengthen AI model testing and evaluation.
Anthropic is an AI safety and research company that's working to build reliable, interpretable, and steerable AI systems.
Focuses on ContextWeave: A Real-World Workflow Benchmark.
Focuses on NSF-HRPT: Neural Semantic Field meets Hierarchical Risk Perception Tree for Safety-Critical Scenario Assessment.
Focuses on A Security-Oriented Lifecycle Model for Large Language Model Systems.