On the Impossibility of Unbiased and Length-Invariant Policy Optimization with Outcome Re…
Explorar
Noticias de IA
22115 elementos — filtrados, clasificados y sin duplicados
Running a 28.9M parameter LLM on an $8 microcontroller
When AI hallucinates aliens: Some think NASA should be nervous
Two Views, One Voice: Evidence-Grounded Conversational Music Recommendation
Evaluating and Mitigating the Misguidance Effect of Buggy Code in LLM-Generated Unit Tests
Same Predictions, Different Reasons: The Effect of Quantization on Model Explanations
PatiGonit22K: A Comprehensive Dataset for Solving Complex Bengali MWPs
Team uses AlphaFold AI to redesign gene-editing proteins to make them safer
OrchNAS: Orchestrated Neural Architecture Search Service for Personalised Federated Edge …
Synthetic Speech, Real Signal: Paralinguistic Preservation and Cross-Lingual Augmentation…
Visual Relocalization from Sparse Views in Aliased and Low-Texture Environments via Novel…
WaveformQA: Benchmarking LLM Temporal Reasoning on Digital Waveforms
Telco-GAIA: Bilingual Benchmark for Agents in Telecom Domain
ElasticTTT: Prior-Preserving Test-Time Tuning for Video Editing
Artificial Epanorthosis: Why large language models overuse a classical rhetorical figure,…
Compact Latent Coordination for Autonomous Vehicles at Unsignalized Intersections
Thinkink: 2D Spatial Ink-native Interaction with LLMs
Adaptive Identity Anchoring: Closed-Loop Keyframe Placement for Synthetic Paired Supervis…
Reexamining zero-shot summarization: Empirical investigation of trustworthiness of LLM-su…
ArbiGraph: Arbitrarily Scalable Verifiable Task Graphs for Evaluating Context Management
Cycle-Consistent and Uncertainty-Aware Neural Surrogates for Tokamak Edge Plasmas
NVIDIA-labs OO Agents: Native Python Object-Oriented Agents
Multimodal Learning for Arcing Detection in Pantograph-Catenary Systems
Agentic Context Management: Solving Agent Memory and Cost by Treating Them as Lifecycle a…
Toward Continuous Assurance for the Democratization of AI Agent Creation in Industry
Agentic coding without the cloud: evaluating open-weight large language models on longitu…
AREX: Towards a Recursively Self-Improving Agent for Deep Research
TOUR: A Trajectory-Level Unlearning Benchmark for Offline Reinforcement Learning
GlucoTune: A Unified Framework for Blood Glucose Preprocessing, Forecasting, and Benchmar…
AI-Driven Multi-Hop Relay Selection for Smart Urban NR-V2X Networks via Learning-to-Optim…