Governed Capability Evolution: Lifecycle-Time Compatibility Checking and Rollback for AI-…
Explorar
Noticias de IA
21270 elementos — filtrados, clasificados y sin duplicados
Qiskit QuantumKatas: Adapting Microsoft's Quantum Computing exercises for LLM evaluation
Governed Metaprogramming for Intelligent Systems: Reclassifying Eval as a Governed Effect
Post-training makes large language models less human-like
JobBench: Aligning Agent Work With Human Will
FineVLA: Fine-Grained Instruction Alignment for Steerable Vision-Language-Action Policies
ScientistOne: Towards Human-Level Autonomous Research via Chain-of-Evidence
Reasoning, Code, or Both? How Large Language Models Handle Variations in Math Questions
Governed Evolution of Agent Runtimes through Executable Operational Cognition
Which Changes Matter? Towards Trustworthy Legal AI via Relevance-Sensitive Evaluation and…
When Eyes Betray AI: Social Gaze Consistency as a Semantic Cue for AI-Generated Image Det…
AGORA: Adapter-Grounded Observation-Action Retention for Inference-Free Prompt Compressio…
MemFail: Stress-Testing Failure Modes of LLM Memory Systems
GENESIS: Harnessing AI Agents for Autonomous 6G RAN Synthesis, Research, and Testing
How to Square Tensor Networks and Circuits Without Squaring Them
A Dataset of Robot-Patient and Doctor-Patient Medical Dialogues for Spoken Language Proce…
What Makes Chain-of-Thought Work at Probe Time? Local Co-occurrence Rather Than Global De…
Multi-Stakeholder LLM Alignment: Decomposing Estimation from Aggregation
TADDLE: A Tool-Augmented Agent for Detecting Deficient LLM-Generated Peer Reviews
LELA: An End-to-end LLM-based Entity Linking Framework with Zero-shot Domain Adaptation
AI-Driven Contribution Evaluation and Conflict Resolution: A Framework & Design for Group…
PaTAS: A Framework for Trust Propagation in Neural Networks Using Subjective Logic
BatteryMFormer: Multi-level Learning for Battery Degradation Trajectory Forecasting
Can Broad Biomedical Knowledge be Contextualized into Scenario-Grounded Propositions?
A Sharper Picture of Generalization in Transformers
Drive-P2D: A Progressive Perception-to-Decision Benchmark for VLMs in Autonomous Driving
CFG-OEC: Classifier Free Guidance with Orthogonal Error Correction
Learning Decentralized LLM Collaboration with Multi-Agent Actor Critic
The Compressive Knowledge Graph Hypothesis: Which Graph Facts Matter for Scientific Hypot…
Maat: The Agentic Legal Research Assistant for Competition Protection