0

Feed

Trending and latest across evals, tools, models, and papers.

Paper4d ago
TTPO: Test-Time Policy Optimization

Recent prominent post-training methods, such as Reinforcement Learning (RL) and On-Policy Self-Distillation (OPSD), have driven rapid progress in mathematical reasoning for large language models, yet their reliance on ground-truth labels precludes test-time training (TTT).

MathematicsPseudo LabellingReinforcement Learning
33stars
0.2/h
Paper4d ago
Puro-2B: Poor Lab's Qwen2-1.5B Trained on RTX 5090 within $5090

Language model pretraining has become almost synonymous with prohibitive cost, placing it out of reach for much of the academic and open-source communities. Although strong open-source efforts already exist, including open-weight models and open-source training recipes, a…

Language Modeling
48stars
0.8/h
Paper4d ago
Generative Semantic Scene Completion

Outdoor LiDAR semantic scene completion (SSC) recovers a dense semantic voxel grid from a scan observing 1% of the target volume, under class imbalance beyond 7,000x. We recast SSC as generative semantic scene completion (GSSC): a single discrete-diffusion formulation in three…

3stars
Paper4d ago
J-Zero: Unified Challenger--Solver--Judge Co-Evolution from Zero Data

Self-evolving language models have recently emerged as a promising path toward superintelligence, with the advantage of reducing the cost of human supervision. While considerable progress has been made in verifiable domains, self-evolution in unverifiable domains remains…

2stars
0.1/h
Model5d ago
Agnes 2.5 Pro Beta
Sapiens AI
$0.15/M151 tok/s
GPQA Diamond91
SciCode48
Humanity's Last Exam (HLE)38
Model5d ago
Qwen3.8-Flash-Next
Alibaba
$0.23/M83 tok/s
GPQA Diamond92
SciCode47
Humanity's Last Exam (HLE)38
Model5d ago
GLM 5.3 Flash
Zai
1.3M$0.24/M50 tok/s
GPQA Diamond91
MedScribe89
TaxEval v276
Paper5d ago
VoiceMem: Streaming Dual-Brain Memory for Real-Time Interaction

Conversational systems, such as duplex speech language models (SLMs), still lack a streaming, accurate, and empathetic memory system as their soul. We introduce VoiceMem, a simple memory architecture with a parallel informational left brain, an emotional right brain, and…

Full-Duplex Spoken DialogueLanguage ModelingRetrieval
444stars
5.0/h
Paper5d ago
V-Rubrics: Visual Faithfulness via Rubric-Based Reinforcement Learning

Vision-language models can produce fluent answers that are insufficiently grounded in the visual evidence: a single unsupported object, chart value, or intermediate inference can undermine an otherwise plausible response.

Instruction FollowingReasoningReinforcement Learning
44stars
0.7/h
Paper5d ago
Code World Model: Coding Agent as World Brain

World models aim to simulate how complex environments evolve under actions and events, yet existing video-based world models primarily learn dynamics from visual observations, which reveal outcomes rather than the underlying knowledge, rules, and mechanisms governing world…

Coding AgentsVideo generationWorld Models
330stars
1.6/h
Paper5d ago
GGSS: Geodesic-Gated Spherical Steering for Inference-Time Debiasing of Generative Vision-Language Models

Generative vision-language models (VLMs) are increasingly used in human-centered settings, yet they can produce demographically biased outputs even when images differ only in controlled attributes such as perceived race or gender.

1stars
Paper5d ago
VBVR-Pro: A Scalable and Verifiable Suite for Native Visual Reasoning

Native visual reasoning treats visual generation as the medium of reasoning itself: visual states (i.e. images and videos) are not merely inputs to be understood or outputs to be rendered, but first-class substrates for problem solving beyond language.

Image generationImage UnderstandingReinforcement LearningVideo generation
24stars
0.0/h
Paper5d ago
Prefix Sliding for efficient test-time scaling

Test-time scaling uses extra test-time compute to improve performance, such as letting language models reason longer when solving a problem. As models keep the entire reasoning trace in memory via full attention, hard tasks that need long thinking can be prohibitively expensive.

ReasoningReinforcement Learning
17stars
0.1/h
Paper5d ago
Skill Issue: Are Skills Language-Invariant in LLMs?

Large language models access knowledge inconsistently across languages, but to what extent do they differ in their skill sets when interacting with different languages? This work quantifies cross-lingual skill inconsistency orthogonally from knowledge and general benchmark…

422stars
Paper5d ago
JIT-Agent: Scaling Harness Intelligence via Just-in-Time Harness Evolution

Agent capability is not determined by the model alone. The agent harness, encompassing memory management, planning strategy, action protocol, and tool/skill orchestration, can dominate the contribution of the underlying foundation model.

AgentsDeep Research Agents
217stars
1.1/h
Paper5d ago
CaSKG: Counterfactual-Causal Skill Graphs for Scalable Agent Skill Retrieval

Reusable skill libraries allow large language model (LLM) agents to reuse procedural knowledge across tasks, but they also turn memory access into a challenging retrieval problem.

0stars
Paper5d ago
Paint What You See: Benchmarking Dexterous Visual Tool Use in Multimodal Agents

Evaluation is shifting from static QA toward agentic settings where models act through external tools. We identify a critical yet underexplored capability within this space - dexterous visual tool use: fine-grained, closed-loop parameterized visual action in which models infer…

0stars
Model6d ago
Granite 4.2 30B
Ibm
$0.28/M78 tok/s
GPQA Diamond64
SciCode37
Humanity's Last Exam (HLE)11
Model6d ago
Granite 4.2 3B
Ibm
$0.05/M234 tok/s
GPQA Diamond56
SciCode25
Humanity's Last Exam (HLE)7
Model6d ago
Granite 4.2 8B
Ibm
$0.11/M152 tok/s
GPQA Diamond63
SciCode30
Humanity's Last Exam (HLE)10
Paper6d ago
GameWAM: A World Action Model for Video Games

Modern video games combine first-person perception, rapid visual changes, persistent world state, and heterogeneous native controls. Existing game agents map visual and task context directly to actions but lack explicit world dynamics modeling, whereas interactive game world…

World Models
28stars
0.1/h
Paper6d ago
LibriBrain100: One Hundred Hours of Broad and Deep MEG Data for Neural Speech Decoding at Scale

We introduce LibriBrain100, a large-scale MEG dataset for speech decoding designed from the ground up for reproducible, standardised evaluation. LibriBrain100 more than doubles the size of the original LibriBrain release, resulting in over 100 hours of high-quality MEG acquired…

14stars
Paper6d ago
DataKernelBench: Can LLMs Optimize Database Queries on GPUs?

GPUs increasingly accelerate database systems, but query-specific peak performance still often relies on hand-written kernels. Existing LLM kernel benchmarks focus on machine learning operators, leaving irregular, heterogeneous, data-movement-heavy database-style operators…

Coding Agents
1stars
Paper6d ago
FrontierChallenge: Evaluating Scientific Workflow Completion

Scientific agents increasingly analyze data, execute code, and produce research artifacts, yet most benchmarks emphasize final answers, isolated programs, or a single domain. We introduce FrontierChallenge, a cross-domain benchmark comprising 300 end-to-end scientific workflows.

Agents
1.3kstars
1.4/h
Paper6d ago
MemUse: Moving Memory Evaluation from Direct QA to Natural Integration in Long-Term Human-AI Conversation

Memory systems for conversational LLMs are conventionally evaluated by direct, fact-seeking questions about prior dialogue (Direct QA): can the model recall fact X from a prior conversation? We tested whether higher Direct QA accuracy correlates with higher user satisfaction in…

0stars
Paper6d ago
WeMM-Embedding: WeChat Multi-Modal Embedding Technical Report

Universal multimodal embeddings are becoming a core component of modern AI systems, enabling heterogeneous content to be represented in a shared space for applications such as retrieval, recommendation, classification, and agentic systems.

Document UnderstandingEmbedding modelsImage UnderstandingOmni models
999stars
2.9/h
Paper6d ago
WarpSAC: Towards the Pinnacle of Scalable Off-policy RL by Rethinking Exploration and Exploitation

Massively parallel simulation changes the data regime in which off-policy reinforcement learning (RL) is trained, challenging stabilizers designed for data-limited replay.

Continuous ControlReinforcement Learning
14stars
Paper6d ago
Recursive Experiential-Working Memory Evolution for Long-Horizon Agent Harnesses

Recursive self-improvement (RSI) remains hard in long-horizon tasks, where growing histories obscure the task state and misalign skill invocation. We introduce Recuris, a recursive Experiential-Working Memory architecture for long-horizon agent harnesses, in which Working Memory…

Agents
122stars
1.1/h
Paper6d ago
StepGuard: Learning Step-Level Guardrails with Scalable Supervision and Safety-Utility Balancing

LLM-based agents can interact with external environments through tool invocation, but this capability also introduces security risks such as file modification, information leakage, and unauthorized actions.

3stars
Paper6d ago
On-policy Distillation with Verifiable Reward

Reinforcement Learning with Verifiable Rewards (RLVR) and on-policy distillation (OPD) have become two widely adopted paradigms for post-training large language models. However, RLVR suffers from sparse task-level feedback, while OPD provides dense token-level guidance but…

ReasoningReinforcement Learning
16stars
0.1/h
Paper6d ago
PonderPounce: A Pretrained MLLM as an Episode Context Engine for Robot Control

Multimodal large language models (MLLMs) can integrate long visual histories, reason under partial observability, and infer behavior from a few examples. Yet vision-language-action (VLA) models generally inherit pretrained representations without using this contextual capacity…

Robotics
3stars
0.1/h
Paper6d ago
Are Android GUI Agents Robust Against Runtime Anomalies? AnTrap: Evaluating Agents in Dynamic Adversarial Environments

GUI agents often encounter dynamic anomalies when deployed on Android devices, from unexpected pop-ups to action misuse, yet existing benchmarks lack systematic evaluation of agent robustness against runtime anomalies.

Computer Use AgentsReinforcement Learning
5stars
Paper7d ago
MARS: Multi-Specialist LLM Relay System for Competitive Programming

Large Language Models excel at code generation, yet competitive programming exposes a persistent failure mode: existing multi-agent pipelines distribute work over generic planner, coder, and debugger roles and delegate the choice of algorithmic technique to the backbone alone.

0stars
Paper7d ago
Autonomous Mathematical Discovery in an Open-World Multi-Agent Environment

We study autonomous mathematical discovery in the Station, an open-world multi-agent environment in which AI agents from different model families pursue a shared research goal without a central coordinator or scripted pipeline.

Agents
176stars
0.3/h
Paper7d ago
CyberFactory: Scaling Cyber Security Capabilities with Instances from the Wild

As large language models (LLMs) continue to advance in coding capabilities, their potential in cybersecurity has drawn increasing research attention, with closed-source LLMs (e.g., Mythos) delivering advanced cybersecurity capabilities.

Coding AgentsQuestion Answering
0stars
Paper7d ago
CaRGo-T: Causal Reasoning Graph-of-Thought improves Multimodal Humor Comprehension

Large-scale vision-language models (VLMs) have demonstrated remarkable versatility across a wide range of multimodal tasks. However, understanding humor remains challenging because humorous content often depends on subtle interactions among entities, events, context, and…

0stars
Paper7d ago
Better Retrieval, Worse Robustness: How Multi-hop RAG Amplifies Upstream ASR Errors

Speech-based applications pass spoken queries through automatic speech recognition (ASR) before any retrieval module, so ASR errors enter the pipeline as a fixed upstream constraint.

Automatic Speech RecognitionQuestion AnsweringRetrieval
3stars
Paper7d ago
Industrial-Instruction: An End-to-End Framework for Building Instruction-Tuning and Benchmark Datasets from Industrial Technical Reports

Industrial technical reports contain high-value knowledge for maintenance, troubleshooting, and product engineering, but their heterogeneous structure (dense prose, specifications, tables) makes them difficult to index and reason over with standard retrieval and QA pipelines,…

Question Answering
0stars
Paper7d ago
The Laws of Context Allocation: Causal Measurement and Closed-Loop Orchestration in Generative Search

As Retrieval-Augmented Generation (RAG) shifts toward diverse portfolio generation, it is stymied by two critical bottlenecks: flawed measurement of evidence utilization, and suboptimal context budget allocation. We resolve both sequentially.

Retrieval
0stars
Paper7d ago
RIBOSPAN: A Long-Context RNA Foundation Model for Versatile RNA Modeling

Full-length RNAs, particularly messenger RNAs, often exceed the context lengths used to pretrain existing RNA foundation models, limiting complete-transcript modeling at single-nucleotide resolution.

Biology
16stars
Paper7d ago
Best Practice Critic Optimization

Group-based reinforcement learning methods such as GRPO for large language models avoid training a critic by sampling multiple responses for each prompt. A reliable critic could instead estimate token-level advantages from one response, but standard critic-based training recipes…

MathematicsReinforcement Learning
10stars
Paper7d ago
ReWorld: An Interactive World Model with Long-Horizon Memory

An interactive world model must follow the user's actions, remember the places it has shown, and stream in real time. The tension is structural: control wants a short horizon, memory wants an unbounded one. ReWorld separates the two during training and bounds them at inference.

Video generationWorld Models
33stars
Paper7d ago
SWE Refactor Bench: Can Coding Agents Complete a Long-Horizon, Whole-Repository Stack Migration?

Modern software systems accumulate technical debt over decades of development, which makes migration expensive and largely manual. As coding agents become increasingly capable at bug fixing, can they autonomously perform such migrations? Existing benchmarks cannot answer this…

Coding Agents
41stars
Paper7d ago
Prime Agent: A Self-Improving RLM Harness

Language models are sequential processors, but long-horizon agency requires external information and computation beyond model weights and active context. Prime Agent is an open-source harness for long-horizon evaluation and coding-agent workflows.

Coding Agents
19kstars
5.8/h
Paper7d ago
Towards a Densing Law for User Representation Learning at Billion-Scale Capacity

User representation learning in real-world industrial scenarios is commonly scaled by increasing user amount, behavioral sequence length and model size. However, existing methods face two challenges: (i) Bottleneck for raw data scaling at billion-scale capacity, as performance…

Click-Through Rate PredictionRecommendation SystemsRetrieval
0stars
Paper7d ago
Agent-G$^2$: Gaussian Guidance for Agentic Reinforcement Learning

Hint-based reinforcement learning addresses reward sparsity in long-horizon agentic tasks by retaining a prefix of an expert trajectory before each rollout, letting the policy explore from a state closer to success.

Reinforcement Learning
39stars
Paper7d ago
Apodex 1.1: Scaling Agentic Intelligence for Complex Work

General-purpose language models can reason and synthesize knowledge, but complex work also requires sustained interaction with files, information sources, and executable code, together with state maintenance, failure recovery, and verifiable delivery.

904stars
Paper7d ago
From Generation to Simulation: How Far Are World Models from Being True Simulators?

With the rapid progress of diffusion models and large-scale video generation, generative world models are increasingly expected to replace traditional simulators, including physics engines, game engines, and reinforcement-learning environments.

Video generationWorld Models
12stars
Paper7d ago
AutoSaddler: Automatic Harness Optimization with Durable Updates from Agent Execution Traces

LLM agents remain unreliable on long-horizon tasks, where small local failures can compound over extended interactions and lead to overall task failure. Although external harnesses can substantially improve robustness, harness design remains a manual and expensive process that…

AgentsCoding Agents
162stars
0.5/h
Paper7d ago
MobilePA-Bench: Benchmarking Mobile Planner Agents on Complex Real-World Tasks

As on-device LLM agents evolve into personal copilots, the mobile operating system has become a key testbed for this paradigm, making rigorous capability evaluation essential.

Reinforcement Learning
3stars
Paper8d ago
ClawProBench: Trace-Aware Evaluation of AI Agents with Runtime Coverage and Frozen Workplace-Style Holdouts

Agent benchmarks often evaluate only final answers even when agents run on stateful runtimes. We argue this under-specifies what is being evaluated: the proper unit is a declared model-plus-runtime configuration whose failures can occur in evidence acquisition, runtime routing,…

Agents
823stars
Paper8d ago
Length-Adaptive Decoding for Masked Diffusion Machine Translation

Machine translation tests masked diffusion language models (dLLMs) because every source token must be rendered faithfully, while fixed canvas decoding must choose target length before denoising.

Language ModelingMachine Translation
2stars
Paper9d ago
GUI-Primitives: Diagnosing Spatial Reasoning Failures in Vision-Language GUI Grounding

Computer-use agents ground natural-language instructions in screenshots to locate interface elements, yet existing benchmarks do not isolate whether models bind relational language to the correct element.

Computer Use Agents
1stars
Model10d ago
DeepSeek V4 Flash Vision (Reasoning, Max Effort)
DeepSeek
131K$0.66/M115 tok/s
GPQA Diamond91
SciCode47
Humanity's Last Exam (HLE)35
Paper10d ago
SecOPD: Mitigating Adaptive Prompt Injections by On-Policy Distillation

Prompt injection is listed as the \#1 threat to AI agents. When an agent accesses external data from websites, files, or emails, an attacker may inject a prompt into the data, saying, &#34;Ignore all prior instructions and perform <an attacker's task>.&#34; To prevent arbitrary…

AgentsAutomatic Speech Recognition
5stars
Paper10d ago
Graph Engineering in the Era of LLM Agents: From Individual Intelligence to System Intelligence

LLMs have evolved from language generators to autonomous agents capable of complex, long-horizon tasks. This evolution has produced paradigms including Prompt Engineering to elicit model capabilities, Context Engineering to manage information access, Harness Engineering to…

Agents
257stars
0.3/h
Paper10d ago
A Modular Agent for Reliable and Auditable Spatial Relation Verification in CT Scans

Reliable spatial understanding is an important prerequisite for future medical vision-language systems that aim to support radiological report generation and structured image understanding.

Image UnderstandingMedical ImagingMedical Visual Question Answering
0stars
Paper10d ago
TLive-Omni: An Omni-Modal Understanding Model for E-Commerce Live Streaming

E-commerce live streaming requires omni-modal understanding of noisy, temporally extended streams, where product facts are distributed across speech, video frames, product images, overlaid text, and user queries.

Automatic Speech RecognitionOmni modelsReinforcement LearningScene Text Recognition
109stars
Paper10d ago
CoToGrasp: Contact-Topology-Conditioned Dexterous Grasp Synthesis via Canonical Workspace Learning

Current dexterous grasp planners primarily optimize for physical stability, focusing on whether an object can be grasped rather than how it should be grasped to support downstream functional tasks.

7stars
Paper11d ago
FlavourBench: Executable Culinary Reward Maps for Language Model Evaluation and Post-Training

Open-ended language-model evaluation often substitutes another model or a small preference panel for a missing answer key. We introduce FlavourBench, which instead compiles dense answer maps from a versioned culinary environment.

5stars
0.0/h
Paper11d ago
Peer-Voted LLM-Agent Stress Tests Find Feed-Induced Lexical Convergence but No Reliable Matched-Exposure Advantage for Distributed Sources

Population-level behavior in large-language-model (LLM) agents cannot be characterized by single-agent benchmarks. We introduce PV-SST, a peer-voted social-platform testbed, and report a separately frozen, preregistered matched-exposure experiment spanning four topics, four…

1stars
Paper11d ago
Towards Quantifying Benchmark Optimization in ASR Models

Public benchmarks are important measures of Automatic Speech Recognition (ASR) model capabilities. However, by nature of being public, there is risk of models being optimized for these benchmarks in ways that do not generalize well to real-world data.

Automatic Speech Recognition
10stars
Paper11d ago
Listening Forward: Next Patch Embedding Prediction Enables Scalable Audio Learners

Self-supervised learning (SSL) has driven substantial progress in audio representation learning, though existing methods have increasingly relied on elaborate pre-training recipes to reach competitive performance.

Audio ClassificationLanguage Modeling
16stars
Paper11d ago
Repo0: Design-Driven Zero-to-All Code Generation

Large language model agents have made substantial progress in code generation, yet most existing systems assume a predefined repository architecture. This assumption does not hold in zero-to-all code generation, where an agent must construct an entire software project directly…

8stars
Paper11d ago
GOAG: Generative and Object-Agnostic Grasp Planner for Dexterous Robotic Manipulation

Multifingered grasping is a crucial robotic skill, but current deep-learning grasp planners often struggle to generalize to new objects because they are trained on limited, object-specific datasets.

Robotics
6stars
Paper11d ago
VGI-Bench: Probing Visual Intelligence in Video Generation Models

Recent studies suggest that video generation models can exhibit certain forms of zero-shot visual reasoning through generated frames. Yet reliable evaluation remains challenging: benchmarks should adopt inputs aligned with the visual priors of current video models, require valid…

Image UnderstandingReasoningVideo generationVideo Understanding
8stars
Paper11d ago
Daedalus-150M: A Convolution-Attention Hybrid Designed for CPU Inference

Small language models are usually built like large ones and then squeezed onto a CPU afterwards. We did the opposite: we fixed the target first, one user, one token at a time, 4-bit weights, ordinary CPU, and chose the architecture to suit it.

Language Modeling
1stars
Paper11d ago
MemTrapBench: Benchmarking Cognitive Traps in LLM Memory Use

Memory has become a key component of large language models, enabling them to retain information and learn from long-term interactions. However, existing memory benchmarks mainly evaluate whether information is correctly extracted, stored, and retrieved, while largely overlooking…

Reasoning
3stars
Paper11d ago
SWE-bench Science: Can Coding Agents Resolve Engineering Tasks in Science?

Software increasingly functions as part of the scientific instrument itself, making failures in scientific code capable of compromising not only program behavior but also the evidence underlying scientific conclusions.

Coding AgentsReinforcement Learning
74stars
Paper11d ago
FlashPrefill V2: Block-Sparse Prefill Attention for Long-Context LLM Serving

Long-context modeling is a pivotal capability for Large Language Models, yet the quadratic complexity of attention remains a critical bottleneck, particularly during the compute-intensive prefilling phase.

Long Context
20stars
Paper11d ago
One Success Isn't Reliability: Thinkingbox, a Sandbox and Benchmark for Agents in Stateful Business Workflows

Recent agent benchmarks increasingly ground evaluation in executable environments, from code repair to web navigation, app APIs, and function calling. Yet completing consequential work beyond code requires more than producing a plausible response or valid tool call: agents must…

Agents
24stars
Paper11d ago
Task-CoEvolve: Efficient Harness Optimization via Adaptive Validation Task Selection

We present a novel approach to efficient LLM harness optimization through adaptive validation task selection. Harness optimization iteratively rewrites the harness code based on validation performance, enabling substantial performance gains without updating the underlying model…

Text classification
11stars
Paper11d ago
EnvHarness: Awakening Static Worlds for Agent Learning

LLM agents learn by interacting with environments, yet these environments are hand-built and static: blind to an agent's weaknesses, and quickly left behind as it improves.

AgentsReinforcement Learning
459stars
2.4/h
Paper11d ago
PolicyGuide: From Guarding One Action to Guiding the Whole Workflow for Policy-Compliant LLM Agents

Customer-service LLM agents must follow organizational policy when acting on a user's behalf. Compliance failures arise from either forbidden actions, such as granting an ineligible change, or omitted procedural requirements, such as identification or confirmation.

Agents
0stars
Paper12d ago
Mechanistic Tomography: Designed Measurement for Control-Oriented Interpretability

Mechanistic interpretability seeks quantities that models do not expose directly: represented states, component effects, interactions, and responses to interventions. Patching, gradients, Hessian-vector products, and subset interventions provide different measurements under…

1stars
Paper12d ago
SPADE: Self-Play in Adaptive Synthetic Executable Environments

Continuous self-improvement requires an ever-expanding pool of self-generated, diverse, adaptive goals. For language agents, existing training environment pools (hand-curated, statically synthesized, or frozen-verifier) keep the goal distribution fixed as the learner scales.

AgentsContinuous ControlReasoningReinforcement Learning
79stars
0.0/h
Paper12d ago
Open-MOPD: Diagnosing and Fixing Capability Imbalance in Multi-Teacher On-Policy Distillation

Multi-teacher on-policy distillation (M-OPD) has emerged as a promising paradigm for consolidating domain-specialized reinforcement learning (RL) experts into a single generalist student via dense, token-level reward supervision.

Instruction FollowingReinforcement Learning
55stars
Paper12d ago
SkillForge: Self-Distilling Agents for Project-Specific Issue Resolution

Large language model (LLM) based agents have demonstrated remarkable proficiency in automated software issue resolution, yet they often struggle to resolve issues in a specific repository because they lack project-specific knowledge.

Reinforcement Learning
5stars
Paper12d ago
SkillGate: Training In-Policy Skill Selection in Long-Horizon Agents

Agent frameworks increasingly package procedural knowledge as skills: instruction files an agent reads on demand, while public libraries now hold thousands of them. Which skill to read has thus become a decision the policy itself makes in the middle of an episode, yet no…

AgentsReinforcement Learning
4stars
Paper12d ago
FACET: Preserving Source Intent and Executable State in Terminal Task Synthesis

Training terminal agents requires scalable executable supervision, yet synthesizing high-quality terminal tasks remains challenging. Each task couples an instruction, an initialized environment, a reference solution, and an executable verifier; if these artifacts are generated…

Coding Agents
119stars
0.2/h
Paper12d ago
Partition the Support, Reconstruct the Residual: Training-Free Sparse Attention for Video Generation and World Models

Training-free block-sparse attention can accelerate video transformers, but row-wise attention concentration does not by itself specify an executable sparse operator. Queries sharing a block route may have poorly overlapping supports, while retained attention mass alone does not…

Video generationWorld Models
15stars
Paper12d ago
FM-Bench: A Benchmark for Long-Horizon Management with Competing Agents

Language model agents now execute bounded tasks reliably. Whether they can sustain effective decision-making over long horizons, where actions have cumulative consequences and the environment responds to their choices, remains largely unmeasured.

Agents
17stars
Model13d ago
GLM 5.3
Zai
1.3M$2.15/M75 tok/s
Creative Story-Writing Benchmark92
GPQA Diamond92
MedScribe89
Paper13d ago
Agentic ESOpt: Fine-Tuning Long-Horizon LLM Agents with Minimal GPU Requirements

Reinforcement Learning (RL) has been promising in single-turn LLM fine-tuning. However, long-horizon agentic reasoning introduces increasingly branching interactions and sparse rewards, exposing several limitations of RL: its heavyweight backpropagation-based training stack…

AgentsReinforcement Learning
58stars
Paper13d ago
AutoResearch: Insight In, Hallucination Out

Autonomous research systems are increasingly capable of executing long research workflows, yet automation alone does not ensure that the resulting process remains scientifically grounded.

Deep Research Agents
192stars
1.6/h
Paper13d ago
HarnessRisk: A Lifecycle-Oriented Benchmark for Agent Harness Safety

Large language models are increasingly deployed through agent harnesses that manage tools, extensions, persistent state, permissions, and external actions. Existing safety benchmarks mainly target individual attack mechanisms or a limited subset of operational settings, making…

AgentsSafety and Grounding
11stars
Paper13d ago
Agent Lightning v1.0: Towards Harnessed Agentic RL

Modern agents operate inside agent harnesses that manage tools, context, and control flow, making the harness a critical part of the agent system. Our original Agent Lightning introduced a disaggregated architecture that connects arbitrary agents to RL training through an LLM…

Coding AgentsInstruction FollowingReinforcement Learning
18kstars
0.5/h
Paper13d ago
SemComp-Bench: Benchmarking Semantic Task Completion in Video Generation

We introduce Semantic Task Completion Video Generation, an outcome-oriented video generation task. Under this formulation, success requires both achievement of the intended outcome and semantic grounding.

Video generation
14stars
Paper13d ago
LEGO-RL: Harness-Native Reinforcement Learning for Coding Agents

Reinforcement learning for coding agents increasingly relies on long-running agent harnesses to manage tool integration, repository contexts, and execution feedback. However, the native execution environments of these harnesses are inherently misaligned with policy-gradient…

Coding AgentsReinforcement Learning
63stars
Paper13d ago
PTXBench: Benchmark and Adapt LLMs for GPU Kernel Optimization with Architecture-specific PTX

We introduce PTXBench, a benchmark for evaluating and adapting large language models (LLMs) to use architecture-specific PTX for GPU kernel optimization. PTXBench measures functional correctness, whether selected target instructions execute at runtime, and speedup over frontier…

14stars
Paper13d ago
ASI-Bench: At the Dawn of Artificial Superintelligence

Artificial superintelligence (ASI) requires AI to move beyond mastering existing knowledge toward exploring the unknown, creating new knowledge, and turning new ideas into verifiable results.

Reinforcement Learning
276stars
0.0/h
Paper13d ago
Co-RL: Unsupervised Reasoning Emerges from Diverse Cohort in Multi-agent RL

Reinforcement learning (RL) has emerged as a powerful approach for improving reasoning in language and vision-language models, yet its strongest successes still depend heavily on ground-truth supervision (e.g., verifiable reward).

Multi-Agent Reinforcement LearningReasoningReinforcement Learning
23stars
Paper14d ago
Benchmarking the Benchmarks: Evaluating Automated Safety Benchmarks for Small Language Models

Small Language Models (SLMs) are increasingly deployed in resource-constrained, privacy-sensitive settings, where safety and bias failures can cause security and societal risks.

Safety and Grounding
0stars
Paper14d ago
Cross-Model Memory Transfer via Target-Side Reader Adaptation

Methods for improving knowledge use in large language models typically fall into two regimes. Non-parametric retrieval offers flexible access to external knowledge, but adds retrieval latency, context overhead, and only shallow integration with the backbone.

Question Answering
1stars
Paper14d ago
The Problem Is the Problem: Towards Scalable Mathematical Discovery

AI systems are increasingly capable of contributing to mathematical research. In research practice, frontier-model reasoning is a limited resource, and expert mathematical review is even more sharply constrained.

Recommendation Systems
11stars
Paper14d ago
TRACE-Bench: Decomposing and Diagnosing Multi-Reference Image Generation

Despite recent advances in unified multimodal models for multi-reference image generation, existing benchmarks remain organized around predefined task types (e.g., &#34;subject composition&#34;), which are ill-suited to this combinatorial setting and lead to fragmented coverage,…

Image generation
0stars
Paper14d ago
$R^3$-Bench: LLMs Struggle with Resource-Rational Reasoning under Shared Budgets

In cognitive science, resource rationality asks how an agent should allocate limited computation to maximize expected value. Most reasoning and agent benchmarks use independent per-task budgets; existing shared-budget studies do not calibrate suite performance against the same…

2stars
Paper14d ago
ClawGym II: Exploring Black-Box RL on Agent Harness

Agent harnesses have substantially improved performance on long-horizon tasks by coordinating agent interactions with the environment. However, reinforcement learning through complex harnesses remains largely unexplored, as scaling such training to long-horizon agent tasks…

Reinforcement Learning
41stars
Paper14d ago
Advancing Open and Reproducible Relational Learning: RelArena-$α$, TabPFN-Rel and RPI

This first release of Prior Labs in relational learning shows our continued commitment to open science. We open-source three pieces of software that we expect to accelerate research in the field towards meaningful real-world impact.

55stars
0.0/h
Paper14d ago
Drive, Pack, Fly: The Travelling Thief Problem with Drone

In collection operations, accumulating payload progressively slows the vehicle, imposing a cumulative penalty on routing efficiency. An onboard drone can offset this penalty by retrieving outlying items, thereby shortening the makespan and increasing operational profit.

Reinforcement Learning
4stars
Paper14d ago
Prior Audit-Repair Context Shifts LLM Verifier Thresholds Toward Leniency

Automated checking pipelines increasingly place one language model as the checker and another (or the same one) as the fixer. We ask whether that wiring changes what the checker reports.

Reasoning
0stars
Paper14d ago
From Sequence to Structure: Relational Uncertainty Propagation for LLM Agents

Reliable uncertainty quantification (UQ) is essential for deploying large language model (LLM) agents in complex interactive environments. Existing UQ methods largely rely on local signals, such as token probabilities, predictive entropy, or per-step confidence, and therefore…

Agents
2stars
Paper15d ago
Large Discovery Models: Empirically-grounded Model-Based Open-Ended Search

Scientific discovery often involves optimising expensive-to-evaluate objectives over vast, structured, and open-ended hypothesis spaces, such as molecules, protein sequences, and computer programs.

24stars
Paper15d ago
UI-Mate: Advancing Open-Weight Foundation GUI Agents with In-Context Demonstrations

Foundation GUI agents can automate complex digital tasks, but deployment is hindered by scarce and biased training data, ambiguous prompts, and unreliable execution. Routine workflows rely on user-specific tools and tacit conventions, so unstated instructions can produce…

Computer Use AgentsReinforcement Learning
85stars
Paper15d ago
Bounded Agents: Delegation Security for Multi-Agent AI Systems

LLM-based agents can act on behalf of a user to access cloud services, call tools, or invoke agents. At session start, the agent's permissions are set but remain static, and each request is evaluated independently, without considering prior actions.

2stars
Paper15d ago
TinyCast: Probabilistic Zero-Shot Forecasting with Computed Periodicity

We introduce TinyCast, an attention-free zero-shot forecaster that emits a predictive distribution from 146,505 parameters, on the premise that at this size the periodic structure of a context is worth computing rather than learning.

3stars
Paper15d ago
Dynamic Multi-Byte Prediction With Hierarchical Language Models

Byte-level hierarchical language models (LMs) have recently emerged as a robust alternative to their popular counterparts that use subword tokenization. However, generating one byte at a time remains a bottleneck for inference speed.

Instruction FollowingMachine TranslationQuestion AnsweringSummarization
5stars
Paper16d ago
Gated Recurrent Transformers: Expressive Depth through Recurrent Modulation

Scaling transformer language models creates an inherent tension between expressivity and memory efficiency. While unique weights across layers preserve functional specialization---from input-grounding to abstract refinement---they incur a substantial memory footprint.

Language Modeling
8stars
0.1/h
Paper16d ago
Prototype-Rectified Iterative Self-supervised Manifold Denoising under Severe Acoustic Shift

Audio-Text Foundation Models (ATMs) fail catastrophically under severe acoustic noise, yet existing adaptation strategies either rely on gradient-based Test-Time Adaptation (TTA), which reinforces noise rather than signal, or on prompt tuning that requires privileged noise…

Audio Classification
2stars
Paper16d ago
VibeWorlding: Can Multimodal Agents Construct 3D Open Worlds End-to-End?

Constructing an interactive 3D open world from a user query is important. However, existing methods are primarily evaluated on idealized, simple queries, making it difficult to systematically analyze and compare how multimodal agents understand user intent, use 3D tools, and…

3D generationContinuous ControlReinforcement Learning
245stars
0.4/h
Paper16d ago
StateM: Reaching 95.3% Raw Accuracy, or a \$15 Frontier Run, on Terminal-Bench 2.1 via Harness Scaling

Long-horizon agents can fail even when their underlying models can solve the constituent steps. They may lose track of mutable state, fail to reactivate lessons from earlier executions, skip known procedures, or stop prematurely.

AgentsCoding Agents
789stars
1.9/h
Paper16d ago
Gathered, Not Admitted: How Attention Brings a Latent Variable into Verbalizable Form

Language models hold latent quantities in a form they can report on, and more of a quantity is present in that form when the task requires reusing it flexibly. What causes a representation to enter that form is open, and the word workspace invites an admission story: a gate that…

1stars
Model17d ago
Qwen3.8 27B
Alibaba
1M$1.13/M46 tok/s
GPQA Diamond91
MedScribe84
TaxEval v271
Paper17d ago
How Do Agents Fail on AutoResearch: End-to-End Diagnostic Evaluation on 100 Real-World Frontier Research Tasks

AI has long assisted scientific research, but the rapid advance of LLMs and agentic scaffolds is reshaping the landscape; a single system can now carry whole-stage research from an initial hypothesis all the way to final published paper, which is a paradigm now referred to as…

23stars
Paper17d ago
Marionette: Predicting World States, Rendering Geometry, Painting Appearance

Interactive game world models typically autoregress visual observations directly in pixel or latent space, forcing structured properties such as pose, geometry, and occlusion to be implicitly maintained by the same generative sequence.

Video generationWorld Models
163stars
Paper17d ago
PACE-Bench: Benchmarking Physics Adaptation via Code Evolution in Dynamic Environments

Self-evolving agents improve future behavior from interaction experience, yet existing evaluations typically optimize under fixed execution conditions and do not test recovery after those conditions change.

Agents
4stars
Paper17d ago
Knowing When to Stop: Bayesian Optimal Stopping for LLM Evaluations

LLM evaluations often use fixed sampling budgets, testing every item the same number of times even after estimates are precise. We introduce optstop, a precision-based adaptive stopping framework that treats evaluation as a sequential measurement problem: keep sampling where…

17stars
0.1/h
Paper17d ago
Can We Defend Against AI-Generated Video Attacks on Real-World Crisis Events? A Systematic Evaluation of Detectors, Generators and Social Dissemination

Recent video generators can fabricate realistic depictions of wars, disasters, public emergencies, and other real-world crises, creating substantial risks of misinformation.

Deepfake Detection
96stars
0.1/h
Paper17d ago
Intern-S2-Mobius: Foundation Model with Decoupled Knowledge and Reasoning

We introduce Mobius-v0, an architecture that comprises a globally shared Memory (FFN) that stores knowledge vectors and multiple Reasoners (Self-Attn) that iteratively achieve compositional reasoning.

Language ModelingReasoning
54stars
Paper17d ago
SimpleOPD: Simple Tokenizer-Agnostic On-Policy Distillation for Long-Context Reasoning

On-policy distillation (OPD) offers a promising way to transfer reasoning capabilities from stronger teacher models, but applying it to long-context reasoning teachers and short-context students introduces practical challenges, including tokenizer mismatch, teacher-student…

MathematicsReasoning
13stars
Paper17d ago
MathForm: Scaling Mathematical Autoformalization with Knowledge Retrieval and Verification-Guided Refinement

Autoformalization is commonly framed as translating natural-language mathematical statements into machine-verifiable formal languages such as Lean 4. However, faithful formalization requires more than translation.

Reinforcement Learning
4stars
Paper17d ago
A Pathway to General-Purpose Scientific AI: Multimodal Comprehension of Scientific Images

Scientific figures and tables encode essential experimental evidence, yet remain difficult for digital libraries and multimodal AI systems to retrieve and interpret. The ALD/E-ImageMiner benchmark and ICDAR 2026 Competition on Information Extraction from Atomic Layer…

Image ClassificationImage UnderstandingQuestion AnsweringSummarization
0stars
Paper17d ago
Demystifying Agent Skills: Why They Work-Until They Don't

Skills have emerged as a practical and effective approach for enhancing LLM agents at inference time through structured packages of knowledge. However, existing evaluations largely measure whether skills improve aggregated task success, leaving a more fundamental question…

7stars
Paper17d ago
ForgeWM: Progressive Causal Training for Few-Step Action-Conditioned Video World Models

Action-conditioned video world models require low-latency causal generation and reliable responses to game-native controls. Although causal distillation enables one- or few-step video synthesis, extending it to interactive world models remains challenging, as discrete keyboard…

Video generationWorld Models
103stars
Paper17d ago
Nanbeige4.2-3B on Apple Silicon: Fixing Deployment Bugs and Decreasing Looped Transformer Memory Overhead

Nanbeige4.2-3B is a 3B-parameter agentic model built around a Looped Transformer (LT) that reuses one stack of layers for a second forward pass, adding effective depth without additional parameters.

Agents
1stars
Paper17d ago
Agentic Transaction: Towards ACID-Compliant Agent Systems

Large language model (LLM) agents are evolving from conversational assistants into autonomous systems that execute long-horizon tasks through reasoning, tool use, code generation, and workspace manipulation.

AgentsReinforcement Learning
8stars
Model18d ago
Gemini 3.7 Flash
Google (Alphabet Inc.)
1.0M$1.5/M345 tok/s
GPQA Diamond90
MedScribe84
TaxEval v275
Paper18d ago
HumanTracker: Towards Comprehensive and Human-Aligned Motion Tracking Benchmark

Humanoid motion tracking is central to teleoperation and whole-body imitation, yet evaluation often disagrees with what people perceive in videos. Kinematic errors average per-frame pose differences but miss the physical artifacts that matter most, particularly unstable support…

Continuous ControlHealthcareRobotics
34stars
Paper18d ago
Who Speaks Matters: Authority-Aware Multi-View RAG over Italian Parliamentary Proceedings

Parliamentary proceedings are a primary record of democratic deliberation, yet their volume and fragmentation make multi-perspective access difficult for citizens, journalists, and researchers.

2stars
Paper18d ago
NaviDC-OCR: Navigating Document Parsing Across Digital and Camera-Captured Documents

Document parsing aims to transform unstructured documents into structured and machine-readable representations. Recent advances in Vision-Language Models (VLMs) have significantly advanced document parsing. However, existing approaches still face two major challenges.

Document Layout AnalysisDocument UnderstandingOCR
168stars
0.2/h
Model19d ago
Qwen3.8 2.4T A95B
Alibaba
1.0M$3/M31 tok/s
GPQA Diamond94
SciCode52
Humanity's Last Exam (HLE)42
Model19d ago
K-EXAONE 2.0 0803
LG AI Research
GPQA Diamond83
SciCode41
Humanity's Last Exam (HLE)19
Model19d ago
A.X-K2
SK Telecom
GPQA Diamond86
SciCode39
Humanity's Last Exam (HLE)30
Model19d ago
Solar Open2 250B
Upstage
GPQA Diamond86
SciCode46
Humanity's Last Exam (HLE)28
Model19d ago
Motif 3
Motif Technologies
GPQA Diamond83
SciCode41
Humanity's Last Exam (HLE)37
Model19d ago
Grok 4.6
xAI
500K$3/M58 tok/s
GPQA Diamond95
MedScribe87
Vibe Code Bench v1.176
Model20d ago
Nemotron 3.5 Lightning
NVIDIA
1M$0.11/M283 tok/s
GPQA Diamond74
SciCode32
Humanity's Last Exam (HLE)11
Model21d ago
Muse Glimmer (high)
Meta Platforms
$0.58/M100 tok/s
GPQA Diamond84
SciCode44
Humanity's Last Exam (HLE)22
Model25d ago
Solar Pro 4
Upstage
524K$0.53/M34 tok/s
GPQA Diamond89
SciCode44
Humanity's Last Exam (HLE)29
Model25d ago
Ling 3.0 Tiny
InclusionAI
197 tok/s
GPQA Diamond73
SciCode24
Humanity's Last Exam (HLE)9
Model26d ago
Muse Spark 1.2
Meta Platforms
1.0M$2/M
GPQA Diamond90
MedScribe90
TaxEval v280
Model27d ago
LFM2.5-2.6B
Liquid AI
GPQA Diamond56
SciCode14
Humanity's Last Exam (HLE)6
Model27d ago
Ling-3.0-flash
InclusionAI
262K$0.11/M371 tok/s
GPQA Diamond86
SciCode41
Humanity's Last Exam (HLE)24
Model28d ago
G9v3-39A5B
AI9Stars
GPQA Diamond81
SciCode34
Humanity's Last Exam (HLE)18
Model28d ago
Qwen3.8 Max
Alibaba
1M$3/M30 tok/s
GPQA Diamond93
MedScribe85
TaxEval v276
Model1mo ago
DeepSeek V4 Flash 0731
DeepSeek
$0.66/M135 tok/s
GPQA Diamond91
SciCode50
Humanity's Last Exam (HLE)39
Model1mo ago
Inkling Small
Thinky
1.0M$0.53/M61 tok/s
GPQA Diamond90
MedScribe84
TaxEval v276
Model1mo ago
Celeris-1
Celeris
$0.33/M1975 tok/s
MMLU-Pro78
GPQA Diamond63
SciCode21
Model1mo ago
Claude Opus 5
Anthropic
1M$10/M54 tok/s
ProofBench99
Creative Story-Writing Benchmark96
MedScribe91
Model1mo ago
Agnes 2.5 Pro Alpha
Sapiens AI
$0.56/M130 tok/s
GPQA Diamond88
SciCode42
Humanity's Last Exam (HLE)34
Model1mo ago
G9v3-3B
AI9Stars
GPQA Diamond44
SciCode18
Humanity's Last Exam (HLE)5
Model1mo ago
Gemini 3.5 Flash Lite
Google (Alphabet Inc.)
1.0M$0.85/M360 tok/s
GPQA Diamond84
TaxEval v273
MedScribe71
Model1mo ago
Gemini 3.6 Flash
Google (Alphabet Inc.)
1.0M$1.5/M175 tok/s
GPQA Diamond93
MedScribe80
TaxEval v275
Model1mo ago
Motif 3 (Beta)
Motif Technologies
GPQA Diamond87
SciCode44
Humanity's Last Exam (HLE)40
Model1mo ago
Kimi K3 (low)
Kimi
$6/M40 tok/s
GPQA Diamond84
SciCode51
Humanity's Last Exam (HLE)25
Model1mo ago
Kimi K3
Moonshot AI
1.0M$6/M36 tok/s
GPQA Diamond94
Creative Story-Writing Benchmark88
MedScribe88
Model1mo ago
Inkling
Thinky
1.0M$1.76/M70 tok/s
GPQA Diamond87
MedScribe85
TaxEval v275
Model1mo ago
Muse Spark 1.1
Meta Platforms
1.0M$2/M
GPQA Diamond90
MedScribe89
OSWorld-Verified81
Model1mo ago
JT-4.1 Flash 236B A21B
China Mobile
GPQA Diamond85
SciCode38
Humanity's Last Exam (HLE)18
Model1mo ago
Grok 4.5
xAI
500K$3/M54 tok/s
GPQA Diamond93
MedScribe87
TaxEval v272
Model1mo ago
Hy3
Tencent
262K$0.24/M94 tok/s
GPQA Diamond90
SciCode48
Humanity's Last Exam (HLE)34
Model2mo ago
Claude Sonnet 5
Anthropic
1M$4/M91 tok/s
LiveBench - Math90
LiveBench - Reasoning87
Vibe Code Bench v1.181
Model2mo ago
LongCat 2.0
LongCat
1.0M$1.3/M44 tok/s
GPQA Diamond78
SciCode35
Humanity's Last Exam (HLE)34
Model2mo ago
GPT-5.6 Luna
OpenAI

The fastest, most cost-efficient tier of OpenAI's GPT-5.6 family, available in limited preview through the OpenAI API and Codex for approved organizations.

1.1M$0.45/M129 tok/sClosed
MedScribe84
Vibe Code Bench v1.177
TaxEval v276
Model2mo ago
GPT-5.6 Terra
OpenAI

The capable, lower-cost tier of OpenAI's GPT-5.6 family, available in limited preview through the OpenAI API and Codex for approved organizations.

1.1M$4.5/M110 tok/sClosed
MedScribe83
TaxEval v276
GPQA Diamond75
Model2mo ago
GPT-5.6 Sol
OpenAI

OpenAI's GPT-5.6 flagship, available in limited preview through the OpenAI API and Codex for approved organizations.

1.1M$8/M84 tok/sClosed
Creative Story-Writing Benchmark87
MedScribe85
ProofBench83
Model2mo ago
GPT-5.5 Instant (June 2026)
OpenAI
$11/M122 tok/s
GPQA Diamond82
SciCode49
Humanity's Last Exam (HLE)20
Model2mo ago
Grok Build 0.1 0616
xAI
$1.25/M
GPQA Diamond90
SciCode50
Humanity's Last Exam (HLE)38
Model2mo ago
GLM 5.2
Zai
256K$2.15/M68 tok/s
τ²-bench (Tau²-bench)99
LiveBench - Math90
MedScribe84
Model2mo ago
Kimi K2.7 Code
Moonshot AI
262K$1.71/M43 tok/s
τ²-bench (Tau²-bench)90
GPQA Diamond90
LiveBench - Reasoning83
Model2mo ago
DiffusionGemma 26B A4B
Google (Alphabet Inc.)
8K
GPQA Diamond67
IFBench59
SciCode34
Model2mo ago
North Mini Code
Cohere
256K108 tok/s
GPQA Diamond76
IFBench58
SciCode38
Model2mo ago
Claude Fable 5
Anthropic

Claude Fable 5 is Anthropic's most capable widely released model, built for the most demanding reasoning and long-horizon agentic work. Shares its base model with the invitation-only Claude Mythos 5.

1M$20/M62 tok/sClosed
τ²-bench (Tau²-bench)99
ProofBench95
LiveBench - Math94
Model2mo ago
Nemotron 3 Ultra
NVIDIA
1M$1.14/M70 tok/s
GPQA Diamond87
τ²-bench (Tau²-bench)83
IFBench81
Model2mo ago
Gemma 4 12B
Google (Alphabet Inc.)
8K$0.15/M142 tok/s
GPQA Diamond66
IFBench45
τ²-bench (Tau²-bench)32
Model3mo ago
Nex-N2-Pro
Nex AGI
262K$1/M134 tok/s
GPQA Diamond89
τ²-bench (Tau²-bench)82
IFBench66
Model3mo ago
MiniMax M3
Minimax
1.0M$0.53/M126 tok/s
GPQA Diamond93
τ²-bench (Tau²-bench)89
MedScribe87
Model3mo ago
Qwen3.7 Plus
Alibaba

Qwen3.7Plus is an AI model from Alibaba.

1M$0.7/M55 tok/sproprietaryClosed
τ²-bench (Tau²-bench)93
GPQA Diamond90
IFBench78
Model3mo ago
Step 3.7 Flash
Stepfun
262K$0.44/M104 tok/s
τ²-bench (Tau²-bench)99
GPQA Diamond81
MathArena69
Model3mo ago
LFM2.5-8B-A1B
Liquid AI
344 tok/s
MATH-50090
IFBench56
GPQA Diamond51
Model3mo ago
Claude Opus 4.8
Anthropic

Claude Opus 4.8 (Adaptive Reasoning, Max Effort) is an AI model from Anthropic.

1M$10/M
τ²-bench (Tau²-bench)94
GPQA Diamond92
MathArena92
Eval3mo ago
Physgym Arena Drhard Public
RL Env

PhysGym Arena DR-hard benchmark for domain-randomized Gym simulator repair

RL Env2 frontier
45
45
45
45
45
Eval3mo ago
Physgym Arena Medley Public
RL Env

PhysGym Arena medley benchmark for achievable medium-hard Gym simulator repair

RL Env2 frontier
100
100
100
100
92
80
Model3mo ago
HyperNova 60B 2605
Multiverse Computing
$0.07/M363 tok/s
GPQA Diamond73
IFBench66
τ²-bench (Tau²-bench)63
Model3mo ago
MiniCPM5-1B
OpenBMB

MiniCPM5-1B (Non-reasoning) is an AI model from OpenBMB.

τ²-bench (Tau²-bench)82
IFBench35
GPQA Diamond27
Model3mo ago
Command A+
Cohere

Command A+ is an AI model from Cohere.

256K245 tok/s
τ²-bench (Tau²-bench)81
GPQA Diamond76
IFBench74
Model3mo ago
Qwen3.7 Max
Alibaba

Qwen3.7Max is an AI model from Alibaba.

1M$3.75/MproprietaryClosed
τ²-bench (Tau²-bench)95
GPQA Diamond92
LiveBench - Math85
Model3mo ago
Gemini 3.5 Flash
Google (Alphabet Inc.)

Gemini 3.5 Flash is an AI model from Google (Alphabet Inc.).

1.0M$3.38/MproprietaryClosed
LiveBench - Math88
LiveBench - Language85
GPQA Diamond83
Model3mo ago
JT-35B-Flash
China Mobile

JT-35B-Flash is an AI model from China Mobile.

τ²-bench (Tau²-bench)99
GPQA Diamond83
IFBench42
Eval3mo ago
Apex Shortlist
RL Env

MathArena Apex Shortlist final-answer evaluation environment

RL Env1 frontier
86
77
60
27
21
Eval3mo ago
Devops Troubleshoot
RL Env

Multi-turn DevOps troubleshooting environment with simulated diagnostic tools

RL Env1 frontier
56
Model3mo ago
MiniCPM-V 4.6 1.3B
OpenBMB

MiniCPM-V 4.6 1.3B is an AI model from OpenBMB.

τ²-bench (Tau²-bench)88
GPQA Diamond31
IFBench27
Eval3mo ago
Teaching Env
RL Env

Evaluates LLM explanations of textbook excerpts across pedagogy dimensions including concept coverage, coherence, prerequisite ordering, and origin...

RL Env1 frontier
76
68
Eval3mo ago
Science Gym Chem
RL Env

Science Sim chemistry compound and reaction screening environment

RL Env1 frontier
88
88
84
66
60
Eval3mo ago
Science Gym Materials
RL Env

Science Sim materials candidate ranking and simulation planning environment

RL Env1 frontier
100
100
75
68
42
Eval3mo ago
Science Gym Bio
RL Env

Science Sim computational biology protein-variant decision environment

RL Env1 frontier
100
100
80
70
45
Model3mo ago
Ring-2.6-1T
InclusionAI

Ring-2.6-1T is an AI model from InclusionAI.

262K$0.85/M127 tok/s
τ²-bench (Tau²-bench)92
GPQA Diamond86
IFBench45
Eval3mo ago
Polars Env
RL Env

Polars DataFrame manipulation environment for training and evaluation

RL Env1 frontier
92
Eval3mo ago
Ar Credit Release V1
RL Env

AR Credit Command Post Evals by Cognida.ai: enterprise mock-ERP credit hold and order release for AR automation agents (structured data only).

RL Env2 frontier
54
50
47
34
Model3mo ago
GPT-5.5 Instant (May 2026)
OpenAI

GPT-5.5 Instant (May 2026) is an AI model from OpenAI.

$11/MproprietaryClosed
GPQA Diamond85
IFBench71
SciCode50
Model4mo ago
Grok 4.3
xAI

Grok 4.3 is an AI model from xAI.

1M$1.56/MproprietaryClosed
LiveBench - Math84
MedScribe74
LiveBench - Language74
Model4mo ago
Mistral Medium 3.5
Mistral AI

Mistral Medium 3.5 is an AI model from Mistral AI.

262K$3/M148 tok/s
τ²-bench (Tau²-bench)94
GPQA Diamond75
IFBench69
Model4mo ago
Granite 4.1 8B
Ibm

granite-4.1-8b is an AI model from Ibm, released with open weights.

131K$0.06/M124 tok/sapache-2.0Open
GPQA Diamond43
IFBench39
τ²-bench (Tau²-bench)28
Model4mo ago
DeepSeek V4 Pro 0423
DeepSeek

DeepSeek's April 2026 next-gen open-weights flagship - 1.6T-total / 49B-active MoE with 1M context and DeepSeek Sparse Attention.

1.0M$0.54/M64 tok/smitOpen
τ²-bench (Tau²-bench)91
LiveBench - Math91
LiveBench - Reasoning83
Model4mo ago
DeepSeek V4 Flash 0423
DeepSeek

DeepSeek V4 Flash is an AI model from DeepSeek, released with open weights.

1.0M$0.17/MmitOpen
τ²-bench (Tau²-bench)94
LiveBench - Math80
MathArena77
Model4mo ago
GPT-5.5
OpenAI

GPT-5.5 is an AI model from OpenAI.

1.1M$11/MproprietaryClosed
Physgym Arena Medley Public100
Crystal Relaxation Rlm100
LiveBench - Math96
Model4mo ago
Hy3 preview
Tencent

Hy3-preview is an AI model from Tencent.

262K$0.1/M
GPQA Diamond73
τ²-bench (Tau²-bench)68
IFBench48
Model4mo ago
MiMo-V2.5-Pro
Xiaomi

mimo-v2.5-pro is an AI model from Xiaomi, released with open weights.

1.1M$0.54/M37 tok/smitOpen
τ²-bench (Tau²-bench)94
GPQA Diamond87
MMLU-Pro85
Model4mo ago
Qwen3.6 27B
Alibaba

Qwen3.6 27B is an AI model from Alibaba.

262K$1.35/M53 tok/s
τ²-bench (Tau²-bench)94
GPQA Diamond83
LiveBench - Math80
Model4mo ago
Kimi K2.6
Moonshot AI

kimi-k2.6 is an AI model from Moonshot AI.

262K$1.71/MModified MITClosed
Physgym Arena Medley Public100
τ²-bench (Tau²-bench)96
GPQA Diamond91
Model4mo ago
Qwen3.6 Max Preview
Alibaba

Qwen3.6 Max Preview is an AI model from Alibaba.

$2.92/MproprietaryClosed
τ²-bench (Tau²-bench)96
GPQA Diamond89
IFBench77
Model4mo ago
Claude Opus 4.7
Anthropic

Claude Opus 4.7 is an AI model from Anthropic.

1M$10/MproprietaryClosed
Physgym Arena Medley Public100
LiveBench - Math93
GPQA Diamond89
Model4mo ago
Muse Spark
Meta Platforms

muse-spark is an AI model from Meta Platforms.

proprietaryClosed
τ²-bench (Tau²-bench)92
GPQA Diamond88
MedScribe86
Model4mo ago
GLM 5.1
Zai

GLM-5.1 is an AI model from Zai, released with open weights.

205K$2.13/MmitOpen
τ²-bench (Tau²-bench)97
MMLU-Pro85
LiveBench - Math85
Model5mo ago
Qwen3.6 Plus
Alibaba

Qwen3.6Plus is an AI model from Alibaba.

1M$1.13/MproprietaryClosed
τ²-bench (Tau²-bench)98
GPQA Diamond88
LiveBench - Math84
Model5mo ago
Gemma 4 31B
Google (Alphabet Inc.)

gemma-4-31b is an AI model from Google (Alphabet Inc.), released with open weights.

36 tok/sapache-2.0Open
GPQA Diamond76
τ²-bench (Tau²-bench)65
IFBench53
Model5mo ago
Gemma 4 26B A4B
Google (Alphabet Inc.)

gemma-4-26b-a4b is an AI model from Google (Alphabet Inc.), released with open weights.

$0.18/Mapache-2.0Open
GPQA Diamond71
IFBench45
τ²-bench (Tau²-bench)40
Model5mo ago
Step 3.5 Flash
Stepfun

step-3.5-flash is an AI model from Stepfun, released with open weights.

262K$0.15/Mapache-2.0Open
τ²-bench (Tau²-bench)87
GPQA Diamond83
Infraresolutionbench76
Model5mo ago
GLM 5V Turbo
Zai

GLM-5v Turbo is an AI model from Zai.

203KproprietaryClosed
τ²-bench (Tau²-bench)99
GPQA Diamond81
LiveBench - Coding74
Model5mo ago
Trinity Large Thinking
Arcee AI

trinity-large-thinking is an AI model from Arcee AI, released with open weights.

262K$0.4/M279 tok/sapache-2.0Open
τ²-bench (Tau²-bench)90
GPQA Diamond75
Infraresolutionbench73
Model5mo ago
MiMo-V2-Omni
Xiaomi

mimo-v2-omni is an AI model from Xiaomi.

262KproprietaryClosed
τ²-bench (Tau²-bench)91
GPQA Diamond83
Infraresolutionbench73
Model5mo ago
MiniMax M2.7
Minimax

minimax-m2.7 is an AI model from Minimax.

197K$0.53/MModified MITClosed
GPQA Diamond87
τ²-bench (Tau²-bench)85
LiveBench - Math81
Model5mo ago
MiMo-V2-Pro
Xiaomi

mimo-v2-pro is an AI model from Xiaomi.

1.0MproprietaryClosed
τ²-bench (Tau²-bench)95
GPQA Diamond87
Infraresolutionbench77
Model5mo ago
GPT-5.4 Nano
OpenAI

GPT-5.4 nano is an AI model from OpenAI.

400K$0.46/MproprietaryClosed
Polars Env92
LiveBench - Math91
LiveBench - Reasoning81
Model5mo ago
GPT-5.4 Mini
OpenAI

GPT-5.4 mini is an AI model from OpenAI.

400K$1.69/MproprietaryClosed
Infraresolutionbench80
LiveBench - Math79
LiveBench - Reasoning72
Model5mo ago
Nemotron 3 Super
NVIDIA

nvidia-nemotron-3-super-120b-a12b is an AI model from NVIDIA.

1M$0.35/M143 tok/sNVIDIA Open ModelClosed
GPQA Diamond80
IFBench71
τ²-bench (Tau²-bench)68
Model5mo ago
Grok 4.20 0309
xAI

Grok 4.20 0309 is an AI model from xAI.

$3/M
GPQA Diamond79
TaxEval v274
τ²-bench (Tau²-bench)70
Model5mo ago
GPT-5.4
OpenAI

GPT-5.4 is an AI model from OpenAI.

1.1M$5.63/MproprietaryClosed
LiveBench - Math94
LiveBench - Reasoning88
GPQA (Full Set)87
Model6mo ago
Gemini 3.1 Flash Lite Preview
Google (Alphabet Inc.)

Gemini 3.1 Flash-Lite is an AI model from Google (Alphabet Inc.).

1.0M$0.56/MproprietaryClosed
GPQA Diamond82
IFBench77
LiveBench - Math74
Model6mo ago
Gemini 3.1 Pro Preview
Google (Alphabet Inc.)

Gemini 3.1 Pro Preview is an AI model from Google (Alphabet Inc.).

1.0M$4.5/M120 tok/s
τ²-bench (Tau²-bench)96
GPQA Diamond94
GPQA (Full Set)93
Model6mo ago
Claude Sonnet 4.6
Anthropic

Claude Sonnet 4.6 is an AI model from Anthropic.

1M$6/M49 tok/sproprietaryClosed
Infraresolutionbench92
Agriculture Qa87
LiveBench - Math87
Model6mo ago
Claude Opus 4.6
Anthropic

Claude Opus 4.6 is an AI model from Anthropic.

1M$10/MproprietaryClosed
LiveBench - Math89
LiveBench - Reasoning89
MedScribe87
Model6mo ago
Qwen3 Coder Next
Alibaba

Qwen3 Coder Next is an AI model from Alibaba.

262K$0.56/M130 tok/sunknownOpen
τ²-bench (Tau²-bench)80
GPQA Diamond74
IFBench35
Model7mo ago
Kimi K2.5
Moonshot AI

Kimi K2.5 is an AI model from Kimi.

262K$1.2/MModified MITClosed
LiveBench - Math85
τ²-bench (Tau²-bench)81
GPQA Diamond79
Model8mo ago
Gemini 3 Flash Preview
Google (Alphabet Inc.)

Gemini 3 Flash Preview (Reasoning) is an AI model from Google (Alphabet Inc.).

1.0M$1.13/M
AIME 2025: Problems from the American Invitational Mathematics Examination97
LiveCodeBench91
GPQA Diamond90
Model8mo ago
GPT-5.2
OpenAI

GPT-5.2 is an AI model from OpenAI.

400K$4.81/MproprietaryClosed
Bb Demo100
LiveBench - Math93
IDE-Bench85
Model9mo ago
DeepSeek V3.2
DeepSeek

DeepSeek V3.2 is an AI model from DeepSeek, released with open weights.

164K$0.32/MmitOpen
LiveBench - Math85
MMLU-Pro84
τ²-bench (Tau²-bench)79
Model9mo ago
Claude Opus 4.5
Anthropic

Claude Opus 4.5 is an AI model from Anthropic.

200K$10/MproprietaryClosed
Nsa Codebreaker100
LiveBench - Math90
MMLU-Pro89
Model9mo ago
Grok 4.1 Fast
xAI

Grok 4.1 Fast is an AI model from xAI.

proprietaryClosed
Medpt100
LiveBench - Math84
LiveBench - Reasoning80
Model9mo ago
Gemini 3 Pro Preview
Google (Alphabet Inc.)

Gemini 3 Pro Preview (low) is an AI model from Google (Alphabet Inc.).

1M$4.5/M
MMLU-Pro90
GPQA Diamond89
AIME 2025: Problems from the American Invitational Mathematics Examination87
Model10mo ago
Claude 4.5 Haiku
Anthropic

Claude 4.5 Haiku is an AI model from Anthropic.

200K$2/M93 tok/s
Mlebench100
DABstep90
MedScribe85
Eval10mo ago
AIME 2025: Problems from the American Invitational Mathematics Examination
Mathematical Association of America

A benchmark for evaluating AI's ability to solve challenging mathematics problems from the 2025 AIME - a prestigious high school mathematics competition.

ActiveMathematics4 frontier
99
97
97
96
96
95
Model11mo ago
Claude Sonnet 4.5
Anthropic

anthropic/claude-sonnet-4.5 is an AI model.

1M$6/MProprietaryClosed
IDE-Bench88
MMLU-Pro86
MedScribe85
Model11mo ago
Gemini 2.5 Flash Preview (Sep '25)
Google (Alphabet Inc.)

Gemini 2.5 Flash Preview (Sep '25) (Non-reasoning) is an AI model from Google (Alphabet Inc.).

1MproprietaryClosed
MMLU-Pro84
MedScribe78
GPQA Diamond77
Model11mo ago
Gemini 2.5 Flash Lite Preview 09-2025
Google (Alphabet Inc.)

Gemini 2.5 Flash-Lite Preview (Sep '25) is an AI model from Google (Alphabet Inc.).

1.0M$0.17/M
MMLU-Pro80
MedScribe76
TaxEval v266
Model11mo ago
Grok 4 Fast
xAI

Grok 4 Fast is an AI model from xAI.

$0.28/MproprietaryClosed
MedScribe82
TaxEval v276
MMLU-Pro73
Model1y ago
GPT-5
OpenAI

OpenAI's August 2025 unified frontier model that auto-routes between a fast model and a deeper "thinking" variant.

400KProprietaryClosed
Gutenberg Env100
MATH100
AIME 2024: Problems from the American Invitational Mathematics Examination95
Model1y ago
GPT-5 Mini
OpenAI

GPT-5 mini is an AI model from OpenAI.

400K$0.69/MproprietaryClosed
Gpu Puzzles Modal100
Csv Qa100
DABstep100
Model1y ago
Claude 4.1 Opus
Anthropic

Claude 4.1 Opus is an AI model from Anthropic.

200K$30/MproprietaryClosed
MMLU-Pro88
GPQA Diamond81
AIME 2025: Problems from the American Invitational Mathematics Examination80
Model1y ago
Qwen3 30B A3B Instruct 2507
Alibaba

Qwen3.30B A3b Instruct 2507 is an AI model from Alibaba, released with open weights.

262K$0.35/Mapache-2.0Open
Science Gym Materials100
Science Gym Bio100
MATH-50098
Eval1y ago
τ²-bench (Tau²-bench)
Sierra

Sierra's dual-control extension of τ-bench - now the user is also an LLM and both agents share access to the same tool-driven environment.

ActiveTool CallingMulti Turn DialogPlanning
99
99
99
99
99
99
Model1y ago
Gemini 2.5 Pro
Google (Alphabet Inc.)

Gemini 2.5 Pro is an AI model from Google (Alphabet Inc.).

1.0M$3.44/MproprietaryClosed
MATH-50097
AIME 2024: Problems from the American Invitational Mathematics Examination89
AIME 2025: Problems from the American Invitational Mathematics Examination88
Model1y ago
Claude 4 Sonnet
Anthropic

Claude 4 Sonnet is an AI model from Anthropic.

200K$6/MproprietaryClosed
MATH-50093
Mini Swe Agent Bench87
Wiki Race84
Model1y ago
Gemini 2.5 Flash
Google (Alphabet Inc.)

Gemini 2.5 Flash is an AI model from Google (Alphabet Inc.).

1.0M$0.85/MproprietaryClosed
Complex Worlds Hack100
MATH-50093
Arena-Hard84
Model1y ago
Qwen3 4B
Alibaba

Qwen3 4B is an AI model from Alibaba.

unknownOpen
Roi Calculator100
Business Valuation94
Irr Calculator92
Model1y ago
Qwen3 0.6B
Alibaba

Qwen3 0.6B is an AI model from Alibaba.

unknownOpen
MATH-50052
Email To Cc Bcc23
GPQA Diamond23
Model1y ago
Qwen3 30B A3B
Alibaba

Qwen3 30B A3B is an AI model from Alibaba.

131K$0.35/Mapache-2.0Open
MATH-50086
Med Agent Bench84
Med Agent Bench84
Model1y ago
DeepSeek V3
DeepSeek

DeepSeek V3 (Dec '24) is an AI model from DeepSeek.

164K$0.49/MunknownOpen
MATH-50089
LiveBench - Instruction Following76
MMLU-Pro75
RL Env1y ago
SWE-Gym
University of California, Berkeley

First open training environment for real-world software-engineering agents - 2,438 Python tasks from 11 repos, each with an executable runtime and a hidden test suite.

RL EnvCode EditingTool CallingDebugging
Eval1y ago
AgentHarm: Harmfulness Potential in AI Agents
UK AI Security Institute (UK AISI)

Assesses whether AI agents might engage in harmful activities by testing their responses to malicious prompts in areas like cybercrime, harassment, and fraud, aiming to ensure safe behavior.

ActiveSafeguards1 frontier
91
Model1y ago
Qwen2.5 Coder Instruct 7B
Alibaba

Qwen2.5 Coder Instruct 7B is an AI model from Alibaba.

unknownOpen
MATH-50066
IFEval61
BIG-Bench Hard (BBH)50
Model1y ago
Qwen2.5-3B-Instruct
Alibaba

Qwen2.5-3B-Instruct is an AI model with 3.0B parameters, released with open weights.

unknownOpen
IFEval65
BIG-Bench Hard (BBH)47
MuSR40
Model1y ago
Qwen2.5 7B Instruct
Alibaba

Qwen2.5-7B-Instruct is an AI model with 7.0B parameters, released with open weights.

33KunknownOpen
IFEval76
BIG-Bench Hard (BBH)54
MATH Level 550
Model1y ago
Qwen2.5-0.5B-Instruct
Alibaba

Qwen2.5-0.5B-Instruct is an AI model with 500M parameters, released with open weights.

unknownOpen
MuSR33
BIG-Bench Hard (BBH)33
IFEval32
Model2y ago
Llama-3.1-8B
Meta Platforms

Llama-3.1-8B is an AI model with 8.0B parameters, released with open weights.

131KunknownOpen
BIG-Bench Hard (BBH)47
MuSR38
MMLU-Pro33
Eval2y ago
τ-bench (tau-bench)
Sierra

Multi-turn customer-service simulation testing whether agents follow domain policies while interacting with a tool-using user simulator.

ActiveTool CallingMulti Turn DialogInstruction Following
63
33
23
Model2y ago
Meta-Llama-3-8B-Instruct
Meta Platforms

Meta-Llama-3-8B-Instruct is an AI model with 8.0B parameters, released with open weights.

8KunknownOpen
IFEval74
BIG-Bench Hard (BBH)50
MMLU-Pro37
RL Env2y ago
BrowserGym
ServiceNow Research

ServiceNow's unified Gym-style framework for web agents - wraps WebArena, MiniWoB, VisualWebArena, WorkArena, AssistantBench, WebLINX, and more under one Playwright-backed interface.

RL EnvPlanningTool CallingBrowser Use
Eval2y ago
GPQA Diamond
New York University

Graduate-level physics, chemistry, and biology multiple-choice questions written by PhDs and verified to be Google-proof.

ActiveScientific ReasoningFactual RecallScience
95
94
94
94
94
93
Eval2y ago
IFEval
Google DeepMind

500 prompts with verifiable instruction-following constraints (word counts, casing, JSON format) checked by deterministic rules - no LLM judge needed.

ActiveInstruction Following2 frontier
90
87
86
84
83
83
Framework2y ago
BenchBuilder
LMArena

LMSYS's automated pipeline for distilling high-quality LLM benchmarks from crowdsourced chat data (e.g. Chatbot Arena, WildChat), producing the Arena-Hard-Auto benchmark.

FrameworkBenchmark Creation
SFT Dataset3y ago
Tülu 3 SFT Mixture
Allen Institute for AI (Ai2)

Allen AI's flagship open SFT mixture combining new persona-driven prompts with curated public data for post-training a frontier-quality instruct model.

SFT DatasetSafetyCode GenerationMath
Preference4y ago
Anthropic HH-RLHF
Anthropic

Anthropic's foundational helpful-and-harmless human preference dataset - the first major public RLHF corpus and a long-time community baseline.

PreferenceSafetyJailbreak ResistanceMulti Turn Dialog
Eval5y ago
Mostly Basic Python Problems (MBPP)
Google Research

974 short crowd-sourced Python tasks with three unit tests each, used alongside HumanEval as a baseline code-generation benchmark.

SaturatedCode GenerationCode2 frontier
100
91
90
86
85
82
RL Env5y ago
ALFWorld
MIT CSAIL

Aligned text-and-3D embodied environment - agents learn household tasks (pick & place, heat, cool, clean) as both TextWorld games and visually-rendered ALFRED scenes.

RL EnvPlanningEmbodiedInstruction Following