Knowing When to Stop: Bayesian Optimal Stopping for LLM Evaluations
Explorar
Noticias de IA
37834 elementos — filtrados, clasificados y sin duplicados
AI Research Preference Models
Residual Dominance as a Structural Account of Last-Item Reliance in Causal Self-Attention…
A Four-Axis Trustworthiness Benchmark for LLM-as-Judge in Principle-Based Regulation
When Personal Memory Has No Single Answer: Evaluating LLM Agents under Irreducible Confli…
The Dynamics of Intelligence Explosions
Wrong but Useful: Trajectory Value Beyond Answer Correctness in Multi-Agent Messages
LLMs Don't Pay for the Jump
Tripwire: Triggering Aligned Refusal via Statistically Certified Safety Neurons
The Past and Future of AI Scientists
AnchorBench: A Multi-Pathway Benchmark for the Anchoring Effect in LLMs
Sensor-Driven Mission Synthesis for UAV/UGV Swarms: A TB-CSPN Coordination Architecture w…
Intern-S2-Mobius: Foundation Model with Decoupled Knowledge and Reasoning
Clearing the Fog: Towards Installing and Refining Proactive Exploration Capabilities in L…
BCMT: Blockwise Causal Memory Transformer
ScienceFlow: A long-horizon agent for ML research, scientific discovery and beyond
Grounding Without Corrective Control: Truth-Tracking Profiles for Large Language Models
Disentangled Shared Representations Improve Morpho-Transcriptomic Integration
From Style Replication to Style Exploration: Enabling Art Style Exploration with Analyze-…
Attributing Preprocessing Invariance in Spectral Foundation Models
MathForm: Scaling Mathematical Autoformalization with Knowledge Retrieval and Verificatio…
Can Language Models Understand mmWave Data? Benchmarking Large Language Models for mmWave…
BiasTrace: Linking Reasoning Behaviours to Biased Outputs in LLMs
Towards Efficient Multimodal and Multilingual Opinion Extraction for STI: A QLoRA-Based F…
Removing Temporal Note Redundancy Improves Multimodal Reinforcement Learning for Medicine
Polaris : Multi Agentic System for Conversational Enterprise Analytics
GBU-Palm: A Multimodal Video Dataset and Benchmark for Palm Presentation Attack Detection
A Hybrid LLM-Based Framework for Automated Security Annotation Generation in Business Pro…
ASSERT: A Measurement Pipeline for GenAI Audits
QuaSAR: Quantization Compensation via Stable Activation-Aware Rank Truncation