Knowing When to Stop: Bayesian Optimal Stopping for LLM Evaluations
Explorar
Noticias de IA
27413 elementos — filtrados, clasificados y sin duplicados
The Dynamics of Intelligence Explosions
SheetCompass: Hierarchical Relation Graphs for Agentic Spreadsheet Reasoning
Twin: Playing an Unknown Game with a Test-Time Digital Twin
Split the Labor: Separating Evidence Interpretation from Decision Aggregation
Handover of In-Context Learning State Across Session Boundaries
Proxy-Validated LLM UX Micro-Simulations: An Artifact-First Protocol for Early-Stage Deci…
Don't Claim Benchmark-Oriented Optimization Improves General Coding Capability -- Diverse…
Does a Language Server Save Tokens for Coding Agents? A Measurement Methodology and Preli…
Think in Latent, Explain in Language: Self-Explainable Latent Reasoning
The Architect: Interactive Visualization of Deep Learning Mathematics Directly in Microso…
Interactive Analysis of Global Explanations using Aggregated Class Activation Maps for Ne…
BCMT: Blockwise Causal Memory Transformer
From Prediction to Intervention: Personalized Meal-Level Glucose Regulation via an LLM Ag…
UltraArUco: A Lightweight Multilingual Library And Framework With Low-Latency Real-Time M…
IterCOMP: Reasoning-aware Adaptive Prompt Compression for Multi-hop Question Answering
Secret-Stego Dissimilarity as a Design Axis: Invertible Coverless Image Steganography wit…
Training-Free Knowledge Transfer Across Model Scales through Activation-Guided Pruning
Fine-Tuning Qwen3-27B for C-to-Rust Code Translation: A Three-Stage Curriculum of Pretrai…
SAGE: Surrogate-gradient Adaptation via Attention-Guided Entropy for Spiking Transformers
CLAIR-Fin: An Adversarial Multi-Agent Framework for Claim-Level Verification and Adaptive…
TeachMateGPT: A Multi-Agent Knowledge-Grounded Framework for Pedagogical Assessment Gener…
Reading Between The Lines: Modeling and Evaluating Behavioral Realism in Legal Simulation
Building AI-Intensive Software with AI: Early Results and a Cautionary Tale on Measuring …
Does ISO-Grounded NFR Specification Improve LLM Code Generation? A Comparison of Rich and…
Data-driven techniques for translational neuroscience and personalized neuro-health
CutClean: Neural Network Pruning for Privacy-Preserving Inference
Do AI chatbots find what experts would? Effects of model, user role, and sample size on s…
PPAPlace: Differentiable Cross-Stage Objectives for Chip Placement Optimization
AdsWorldEngine: A Self-Evolving Conversational Advertising Agent through Orchestrator and…