The Metanym Game: A Self-Contained, Self-Consistent LLM Peer-Community Benchmark for Stru…
Explorar
Noticias de IA
22318 elementos — filtrados, clasificados y sin duplicados
UniRank: Unified Rank Allocation for Low-Rank LLM Compression
Intent-Governed Tool Authorization for AI Agents
AgentDSE: Reasoning-Augmented Architectural Design Space Exploration
Plans Don't Persist: Why Context Management Is Load Bearing for LLM Agents
Text-to-Image Generative AI for Modeling and Simulation: Methods, Opportunities, and Appl…
Detecting Satellites in Radio-Frequency Data via Semi-Supervised Learning
Not All Claims Are Equally Risky: FACTOR for Adaptive Verification in Factual Long-Form G…
Generating Public Health Responses using Survey-Augmented Large Language Models
A Matter of Time: Towards a General Theory of Agency
DEMM-Bench: A Cross-Regime Benchmark for Agent-Runtime Governance-Evidence Sufficiency
Human and AI collaboration for pulmonary nodule segmentation
Is Agent Code Less Maintainable Than Human Code?
Scaling Linear Mode Connectivity and Merging to Billion Parameter Pretrained Transformers
CalVerT: Augmenting Agents with Calibrated Verifier Telemetry Improves Action and Learnin…
Learning What Not to Forget: Long-Horizon Agent Memory from a Few Kilobytes of Learning
Interpretable Uncertainty Routing Separating Emotion Ambiguity from Distribution Shift in…
Scene-Level Heterogeneous Physics Simulation with 3D Gaussian Splats
When Confidence Takes the Wrong Path: Diagnosing Retrieval-State Lock-In in RAG
EHR-Complex: Benchmarking Medical Agents for Complex Clinical Reasoning
Cohort Organized Learning: Clustering Through Agreement
Litmus: Zero-Label, Code-Driven Metric Specification for Evaluating AI Systems
CADRE: Stable, Parameter Efficient Adaptation of Medical Vision Language Models with Boun…
Denoising Iterative Self-Correction: Structured Verification Loops for Reliable LLM Reaso…
PrivacyAlign: Contextual Privacy Alignment for LLM Agents
Clinical Term Extraction using Open-Source Small Language Models
TACO: Task-Aware Column Description Generation Using LLMs
POTracker: Optimizing Large Language Models for Standard-Compliant Power Outage Report Ge…
Decodable but Not Faithful: Coupling Natural-Language Rationales to Programmatic Verifiers
When Is an LLM Worth It for Hyperparameter Optimization? A Budget-Matched Study on Tabula…