Efficient SWE Agent Benchmarking via Trajectory-Aware Evaluation
Focuses on Efficient SWE Agent Benchmarking via Trajectory-Aware Evaluation.
Topic archive
Evaluation, inference quality, and model cognition. This page collects the latest briefings that match the topic so readers can follow one area without scanning the full feed.
Indexed briefings
619
Latest source-linked updates, ordered newest first.
Latest
Focuses on Efficient SWE Agent Benchmarking via Trajectory-Aware Evaluation.
Focuses on TAG-Bench: Benchmarking Temporal Audio Grounding in Large Audio Language Models.
Focuses on EvoSCM: Scientific Belief Revision Through Causal Model Evolution and Experimentation.
Focuses on Context-Aware Interleaved Batching for WhisperX.
Focuses on Aspire: Can Models Self-Evolve from Vague Goals?.
Focuses on Defense-as-Skill: Evolving Runtime Guard Skill for Skill-Augmented Agents.
Focuses on S3Gym: Can LLMs Turn Self-Testing and Self-Judging into Self-Improvement?.
Focuses on Parsing the Stream: A Live Trace Model for Long-Horizon Agents and Their Observers.
Focuses on Token-Efficient Data Reasoning Agents via Adaptive Structuring of Unstructured Data.
Focuses on Scaling Large Reasoning Models beyond Human Supervision: A Path toward Superintelligence.
Focuses on EDGE: Error Dependency Graph-Guided Multi-Error Attribution in Multi-Agent LLM Systems.
Focuses on LEAP: Likelihood Elicitation and Aggregation for LLM-based Probabilistic Forecasting.
Focuses on LLM-Based Agents for Software and Systems Security: Approaches, Applications, and Assessment.
Focuses on Evidence-Bounded Mental Health Reasoning from Heterogeneous Speech Protocols.
Focuses on COVER: Identifiable Evaluation of Coalition Routing.
Focuses on Explore Before Committing: Hypothesis-Guided Search for Deep Research Agents.
Focuses on A Human-in-the-Loop Autonomous Agent for Industry Time Series Forecasting.
Focuses on Learning to Use Tools: Reinforcement Learning for Tool-Integrated Mathematical Reasoning.
Focuses on When Verified Source Becomes Attack Input: Defending Smart Contracts Against LLM-Based Vulnerability Scanning.
Focuses on RetailAgent: Structured Adverse Timing in Self-Conditioned Multimodal LLM Trading Agents.
Focuses on PersonaForge: Realistic Multi-Turn User Simulation for Agentic Systems.
Focuses on TRIPPULSE: Multi-Agent Travel Planning with Review-Grounded Reasoning.
Focuses on Continuous Autonomous Refactoring: A Research Roadmap for AI-Driven Code Quality Maintenance.
Focuses on EvoUndo: Recoverability-Constrained Self-Evolution for LLM Agent Harnesses.
Focuses on AGENT-O: A Semantic Agent Card Framework for Interoperable and Governed Healthcare AI Agents.
Focuses on LoopArena: Benchmarking Models as Runtime Controllers for Loop Engineering.
Focuses on STEGNav: Spatio-Temporal Event Graph Reasoning for Multimodal Lifelong Object Navigation.
Focuses on Parser States Already Know: Structure-Conditioned KV Persistence for Structured Generation.
Focuses on CoCoBench: A Cooperative Coordination Benchmark for Embodied Multi-Agent Task Planning.
Focuses on Finding Where the Buck Stops: An Automated Failure Attribution-Based Reflection Framework for Multi-Agent Collaboration.
Focuses on Beyond Task-Only Matching: Personalized Skill Routing with Counterfactual Evaluation.
Focuses on Schwarz: Solver-Aware Agentic Program Verification.
Focuses on Adaptive Strategy Generation for Boundary Value Exploration Beyond Numeric Inputs.
Focuses on MuSP-Bench: Advanced Multimodal Benchmarking of Music Understanding across Score and Performance.
Focuses on Benchmarking large language model agent societies against human behavioural distributions.
Focuses on RedEvoAgent: Automatic Red-Teaming Agent with Experience-Driven Skill Evolution.