Identifying and Understanding Human Values in Text: A Tailorable LLM-based Architecture
Explorar
Noticias de IA
37834 elementos — filtrados, clasificados y sin duplicados
Sense Representations Are Inducible Interfaces
EAPO: Entropy-Driven Adaptive Positive-Negative Sample Weighting for Policy Optimization …
On the Origin of Synthetic Information by Means of Steganographic Inheritance
DynaSchedBench: Calibrated Dynamic Scheduling Benchmarks and Observability Paradox in LLM…
Discovery Agents for Real-Time Analytics: Toward Proactive Insight Systems
Agyn: An Open-Source Platform for AI Agents with Scalable On-Demand Execution, Agent Defi…
Laguna M.1/XS.2 Technical Report
ReflexGrad: Within-Episode Failure Recovery in LLM Agents via Progress-Gated Dual-Process…
Reasoning and Planning with Dynamically Changing Norms
Intelligence as Managed Autonomy: Failure, Escalation, and Governance for Agentic AI Syst…
DiagramBank: A Quality-Audited Dataset of Scientific Schematic Diagrams with Multi-Level …
Do Clinical Models Change Treatment Decisions?
DeepSciVerify: Verifying Scientific Claim--Citation Alignment via LLM-Driven Evidence Esc…
Probability-Entropy Calibration: An Elastic Indicator for Adaptive Fine-tuning
Path Channels and Plan Extension Kernels: a Mechanistic Description of Planning in a Soko…
STFlow: Data-Coupled Flow Matching for Geometric Trajectory Simulation
PEAM: Parametric Embodied Agent Memory through Contrastive Internalization of Experience …
Diagnosing Live Within-Policy Instruction Conflicts in LLM Agents with Witnessed Resoluti…
A Query Engine for the Agents
GradientStabilizer:Fix the Norm, Not the Gradient
Adapting the Interface, Not the Model: Runtime Harness Adaptation for Deterministic LLM A…
A Fixed-Budget, Cluster-Aware Standard for LLM-as-a-Judge Evaluation: A Multi-Hop RAG Str…
EgoBench: An Interactive Egocentric Multimodal Benchmark for Tool-Using Agents
TCP-MCP: Landscape-Guided Co-Evolution of Prompts and Communication Topologies for Multi-…
Text2Model: Modeling Copilots for Text-to-Model Translation
DIG to Heal: Scaling General-purpose Agent Collaboration via Explainable Dynamic Decision…
SuiChat-CN: Benchmarking Contextual Suicide Risk Assessment in Chinese Group Chats
STAB: Specification-driven Testing for Algorithmic Bottlenecks
An Empirical Audit of k-NAF Budget Accounting for Anchored Decoding