Efficient SWE Agent Benchmarking via Trajectory-Aware Evaluation
Focuses on Efficient SWE Agent Benchmarking via Trajectory-Aware Evaluation.
Source archive
Fresh research papers across agents, reasoning, multimodal systems, and AI tooling. This source page gives readers and crawlers a stable route for the latest arXiv coverage.
Indexed briefings
2966
Latest source-linked updates, ordered newest first.
Latest
Focuses on Efficient SWE Agent Benchmarking via Trajectory-Aware Evaluation.
Focuses on CordisBench: Can Language Models Reason About Component Lifecycles in Dynamic Agent Harnesses?.
Focuses on The Rise of Verbal Reinforcement Learning.
Focuses on Mechanism Design for Alignment and Control.
Focuses on Designing Proactive Thought Partners for Writing.
Focuses on Selective Agent Guidance via Entropy: Learning Autonomous Policies from Imperfect VLM Teachers.
Focuses on Retrieved but not ranked: surface-form bias in structural retrieval, from mathematics to agent trajectories.
Focuses on NashDreamer: Model-Based Reinforcement Learning for Zero-Sum Imperfect-Information Games.
Focuses on TAG-Bench: Benchmarking Temporal Audio Grounding in Large Audio Language Models.
Focuses on EvoSCM: Scientific Belief Revision Through Causal Model Evolution and Experimentation.
Focuses on Context-Aware Interleaved Batching for WhisperX.
Focuses on When Guardrails Look Effective: Construct Validity Failures in LLM Agent Commerce Evaluation.
Focuses on DIASENTINEL: An Auditable Multi-Agent System for Guideline-Grounded Diabetes Risk Screening.
Focuses on GlossoGen: Emergent Language in Complex Multi-Agent LLM Interactions.
Focuses on Aspire: Can Models Self-Evolve from Vague Goals?.
Focuses on Defense-as-Skill: Evolving Runtime Guard Skill for Skill-Augmented Agents.
Focuses on DreamX-Creator: Democratizing Native Audio-Video Generation at 2K Resolution.
Focuses on Harness-of-Harness: Multi-Day Autonomous Software Development with Continual Improvement.
Focuses on S3Gym: Can LLMs Turn Self-Testing and Self-Judging into Self-Improvement?.
Focuses on Parsing the Stream: A Live Trace Model for Long-Horizon Agents and Their Observers.
Focuses on Token-Efficient Data Reasoning Agents via Adaptive Structuring of Unstructured Data.
Focuses on HarnessDev: Can LLMs Create and Evolve Their Own Agent Harness?.
Focuses on Reconciling Process Supervision with Outcome-Based Credit in Agentic Policy Optimization.
Focuses on TRIAGE: Three-level Routing and Intelligent Agent Guidance for Efficient Execution.
Focuses on Learning to Evaluate Before Improving: Automatic Rubric Induction for Automatic Research Agents.
Focuses on Provably Safe Sim-to-Real Transfer.
Focuses on Scaling Large Reasoning Models beyond Human Supervision: A Path toward Superintelligence.
Focuses on EdiTikZ: Scientific Figure Editing from Revision Trajectories.
Focuses on Measure Before You Manage: Evaluating Agent Working Memory in Coding Agents.
Focuses on Evaluating Multimodal LLMs as Generalist Vision-Language-Action Agents for Drone Control: Commanding, Approaching, Tracking and Searching.
Focuses on Logos: An Agent Harness on a Cross-Process Bus.
Focuses on When Robots Mishear Us: Mapping the Safety Risks of Voice-Controlled Embodied AI.
Focuses on Language-Statistical Analysis of Neural Audio Codec Tokens Across Architectures, Corpora, and Noise Conditions.
Focuses on Autonomous robotic bridging using distributed swarm control without inter-agent communication.
Focuses on Phoneme- and Word-Level Metrics Using Self-Supervised Speech Representations for Forced Alignment Evaluation.
Focuses on When Does Predictor-Based RL Align with Human Perception? A Study of Subjective Rewards in Codec-Based Speech Language Models.