-
Trace2Skill: Distill Trajectory-Local Lessons into Transferable Agent Skills
Paper • 2603.25158 • Published • 49 -
SkillsBench: Benchmarking How Well Agent Skills Work Across Diverse Tasks
Paper • 2602.12670 • Published • 59 -
SkillProbe: Security Auditing for Emerging Agent Skill Marketplaces via Multi-Agent Collaboration
Paper • 2603.21019 • Published
Collections
Discover the best community collections!
Collections including paper arxiv:2602.12670
-
Endless Terminals: Scaling RL Environments for Terminal Agents
Paper • 2601.16443 • Published • 18 -
Linear representations in language models can change dramatically over a conversation
Paper • 2601.20834 • Published • 21 -
Scaling Embeddings Outperforms Scaling Experts in Language Models
Paper • 2601.21204 • Published • 102 -
Teaching Models to Teach Themselves: Reasoning at the Edge of Learnability
Paper • 2601.18778 • Published • 42
-
ReportBench: Evaluating Deep Research Agents via Academic Survey Tasks
Paper • 2508.15804 • Published • 15 -
Behavioral Fingerprinting of Large Language Models
Paper • 2509.04504 • Published • 6 -
Statistical Methods in Generative AI
Paper • 2509.07054 • Published • 11 -
CLUE: Non-parametric Verification from Experience via Hidden-State Clustering
Paper • 2510.01591 • Published • 28
-
Exploring Reasoning Reward Model for Agents
Paper • 2601.22154 • Published • 23 -
Group-Evolving Agents: Open-Ended Self-Improvement via Experience Sharing
Paper • 2602.04837 • Published • 9 -
Agent Skills: A Data-Driven Analysis of Claude Skills for Extending Large Language Model Functionality
Paper • 2602.08004 • Published • 5 -
SEAD: Self-Evolving Agent for Multi-Turn Service Dialogue
Paper • 2602.03548 • Published • 4
-
Benchmark^2: Systematic Evaluation of LLM Benchmarks
Paper • 2601.03986 • Published • 34 -
BabyVision: Visual Reasoning Beyond Language
Paper • 2601.06521 • Published • 201 -
Lost in the Noise: How Reasoning Models Fail with Contextual Distractors
Paper • 2601.07226 • Published • 33 -
CAR-bench: Evaluating the Consistency and Limit-Awareness of LLM Agents under Real-World Uncertainty
Paper • 2601.22027 • Published • 85
-
Describe What You See with Multimodal Large Language Models to Enhance Video Recommendations
Paper • 2508.09789 • Published • 5 -
MM-BrowseComp: A Comprehensive Benchmark for Multimodal Browsing Agents
Paper • 2508.13186 • Published • 20 -
ZARA: Zero-shot Motion Time-Series Analysis via Knowledge and Retrieval Driven LLM Agents
Paper • 2508.04038 • Published • 1 -
Prompt Orchestration Markup Language
Paper • 2508.13948 • Published • 48
-
Trace2Skill: Distill Trajectory-Local Lessons into Transferable Agent Skills
Paper • 2603.25158 • Published • 49 -
SkillsBench: Benchmarking How Well Agent Skills Work Across Diverse Tasks
Paper • 2602.12670 • Published • 59 -
SkillProbe: Security Auditing for Emerging Agent Skill Marketplaces via Multi-Agent Collaboration
Paper • 2603.21019 • Published
-
Exploring Reasoning Reward Model for Agents
Paper • 2601.22154 • Published • 23 -
Group-Evolving Agents: Open-Ended Self-Improvement via Experience Sharing
Paper • 2602.04837 • Published • 9 -
Agent Skills: A Data-Driven Analysis of Claude Skills for Extending Large Language Model Functionality
Paper • 2602.08004 • Published • 5 -
SEAD: Self-Evolving Agent for Multi-Turn Service Dialogue
Paper • 2602.03548 • Published • 4
-
Endless Terminals: Scaling RL Environments for Terminal Agents
Paper • 2601.16443 • Published • 18 -
Linear representations in language models can change dramatically over a conversation
Paper • 2601.20834 • Published • 21 -
Scaling Embeddings Outperforms Scaling Experts in Language Models
Paper • 2601.21204 • Published • 102 -
Teaching Models to Teach Themselves: Reasoning at the Edge of Learnability
Paper • 2601.18778 • Published • 42
-
Benchmark^2: Systematic Evaluation of LLM Benchmarks
Paper • 2601.03986 • Published • 34 -
BabyVision: Visual Reasoning Beyond Language
Paper • 2601.06521 • Published • 201 -
Lost in the Noise: How Reasoning Models Fail with Contextual Distractors
Paper • 2601.07226 • Published • 33 -
CAR-bench: Evaluating the Consistency and Limit-Awareness of LLM Agents under Real-World Uncertainty
Paper • 2601.22027 • Published • 85
-
ReportBench: Evaluating Deep Research Agents via Academic Survey Tasks
Paper • 2508.15804 • Published • 15 -
Behavioral Fingerprinting of Large Language Models
Paper • 2509.04504 • Published • 6 -
Statistical Methods in Generative AI
Paper • 2509.07054 • Published • 11 -
CLUE: Non-parametric Verification from Experience via Hidden-State Clustering
Paper • 2510.01591 • Published • 28
-
Describe What You See with Multimodal Large Language Models to Enhance Video Recommendations
Paper • 2508.09789 • Published • 5 -
MM-BrowseComp: A Comprehensive Benchmark for Multimodal Browsing Agents
Paper • 2508.13186 • Published • 20 -
ZARA: Zero-shot Motion Time-Series Analysis via Knowledge and Retrieval Driven LLM Agents
Paper • 2508.04038 • Published • 1 -
Prompt Orchestration Markup Language
Paper • 2508.13948 • Published • 48