BA Agent RL Environment and Benchmark
RL env & benchmark for enterprise BA agents
Healthcare AI, Medical AI, Vision AI, Audio AI, AI Safety, Agentic Systems, Multilingual AI, Physical AI, Model Fine-tuning, Safety Evaluation, Data Annotation, Localization, Natural Language Processing, Benchmarking, AI Evaluation
BCoughBench: Benchmarking Respiratory Acoustic Foundation Models Under Body-Coupled Wearable Sensor Conditions
Beyond Classification: A Cough Regression Benchmark for Respiratory Acoustic Foundation Models
RL env & benchmark for enterprise BA agents
Simulate and evaluate personal assistant actions on a virtual iPhone
Explore iOS assistant benchmark tasks and view results
Generate personalized ads with instant quality scoring
Cached replays of 140 agent-to-agent negotiation rollouts
Document-work benchmark for healthcare persona
A three-tier rubric for the Finance sector of GDPval
RL environment for sales & revenue-ops agents
RL environment & benchmark for clinical EHR agents
Co-evolutionary adversarial training demo (DA vs CA)
Explore RL tasks and view model trajectories
Interactive demo for the MedMosaic medical-audio benchmark