🔥
热榜聚合
· arXiv AI
← 返回首页
arXiv AI实时榜单
更新于 09-11 17:20
共 30 条
收录至热榜聚合 · 全网 30+ 平台
1
OpenDiscoveryTrace: Process Traces for Evaluating AI Scientist Workflows
OpenDiscoveryTrace:用于评估AI科学家工作流的流程追踪
2
Adaptive Entangled Game Modules in Artificial General Intelligence
通用人工智能中的自适应纠缠博弈模块
3
Subagents vs Agent Skills: Executing Reusable Knowledge for Long-Horizon Agentic Tasks
子代理与代理技能:为长周期代理任务执行可复用知识
4
Gradland: On Phenomenal Experience, Differentiated Across Many Dimensions
Gradland:论现象体验的多维度分化
5
An Autonomous GeoAI Agent for Arctic Eco-Navigation
面向北极生态导航的自主GeoAI代理
6
The Menu Is an Execution Prior: State-Path Tool Menus for Online Agents
菜单即执行优先级:面向在线代理的状态路径工具菜单
7
Decision-Focused Active Learning for Scale-Aware Critical-Materials Recovery
面向规模感知关键材料回收的决策聚焦主动学习
8
Valerant: An Automatic Navigable Game Map Generator via Action-Conditioned World Model Exploration
Valerant:通过动作条件世界模型探索的自动可导航游戏地图生成器
9
XAI-Arena: Can LLMs Assess the Quality of XAI Explanations?
XAI-Arena:LLM能否评估XAI解释的质量?
10
Do Agents Know When They Succeed? Calibrating Agent Confidence from Internal Representations
智能体知道自己何时成功吗?从内部表征校准智能体置信度
11
ContractEval: Query-Conditioned Execution Matching for Procedural Instruction Conformance
ContractEval:面向程序化指令符合性的查询条件化执行匹配
12
Multi-Agent Agentic Graph Learning via Structural Signatures
基于结构签名的多智能体图学习
13
CityPlanner: A Sandbox Agent for Executable Urban Planning
CityPlanner:面向可执行城市规划的沙盒智能体
14
A Function-Space Approach to the Statistical Mechanics of Learning Dynamics
学习动力学统计力学的函数空间方法
15
From State Synchronization to Cognitive Self-Evolution: An Operational Architecture for Cognitive Digital Twins
从状态同步到认知自演化:认知数字孪生的可操作架构
16
Seven Sources of Physical AI Capability Formation
物理AI能力形成的七个来源
17
RobustSGPO: Search-Space Control for Agent Harness Evolution
RobustSGPO:智能体框架演化的搜索空间控制
18
Black-Box Red Teaming of Agentic AI: A Taxonomy-Driven Framework for Automated Risk Discovery
智能体AI的黑盒红队测试:面向自动化风险发现的分类驱动框架
19
RESCUE-BENCH: Towards Relation-Aware Multi-Party Emotional Support Conversation Systems
RESCUE-BENCH:面向关系感知的多方情感支持对话系统
20
PRAGMA: Evaluating Personalized Guidance with Memory Alignment in Lifelong Conversations
PRAGMA:在终身对话中评估记忆对齐的个性化引导
21
Safe to Stop? Risk-Constrained Stopping for Sequential Clinical Diagnosis Agents
安全停止?面向序列临床诊断智能体的风险约束停止
22
Decision Shifts, Lost Label Functionality, and an Inconclusive Grounding Audit in Correctness-Gated Multi-Teacher Distillation
决策偏移、标签功能丧失与正确性门控多教师蒸馏中的不确定性基础审计
23
Which Tokens Should SFT Actually Learn? A Token-Trimming Perspective on Mathematical Reasoning
SFT究竟该学哪些token?数学推理的token裁剪视角
24
Can Artificial Intelligence Support Healthcare and Mental Health Through Early Cyberbullying Detection ? The Impact of Emotion-Aware AI on Proactive Online Safety
人工智能能否通过早期网络欺凌检测支持医疗与心理健康?情感感知AI对主动在线安全的影响
25
LexAgentHallu: A Hierarchical Benchmark for Profiling Hallucinations in Legal Agents
LexAgentHallu:刻画法律智能体幻觉的分层基准
26
Procedural Memory Under Change: Reuse and Interference in Controlled Web Tasks
27
Proof-Carrying Cognition: Closing the Verification Gap with Reality-Settled Reward
28
UnitBoost: Managing Compound LLM Systems with a Merge Operator, Not a Model
29
The Era by Eon Benchmark: A Generated Enterprise Estate with Exact Ground Truth for Benchmarking LLM Agents
30
Shifting Relational Paradigms for Affective Computing: Affective Resonance, Vitality Affects, and Vocal Interaction Fields
← 返回热榜聚合首页
· 数据每日更新 · 热榜聚合 500+ 平台
已经到我的底线了