Evaluating LLM-Based 0-to-1 Software Generation in End-to-End CLI Tool Scenarios
Explorar
Noticias de IA
27413 elementos — filtrados, clasificados y sin duplicados
Decoupled Alignment for Robust Plug-and-Play Adaptation
Soft $Q(\lambda)$: A multi-step off-policy method for entropy regularised reinforcement l…
RLearner-LLM: Balancing Logical Grounding and Fluency in Large Language Models via Hybrid…
MAnchors: Memorization-Based Acceleration of Anchors via Rule Reuse and Transformation
What Is the Minimum Architecture for Prolepsis? Early Irrevocable Commitment Across Tasks…
When Perplexity Lies: Generation-Focused Distillation of Hybrid Sequence Models
GraphDx: A Cost-Aware Knowledge-Enhanced Multi-Agent Framework for Sequential Diagnosis
Energy-based Transport for Amortized Bayesian Inference
Diagnosing Overhead in Dispatch Operations: Cross-architecture Observatory
Scaling Point-in-Time Language Models
Interaction-Aware Whole-Body Control for Compliant Object Transport
ToolSciVer: Multimodal Scientific Claim Verification with Visual Tool Augmented Reinforce…
From Reconstruction to Interpretation: Zero-Setup Multi-Phase Segmentation of X-ray Tomog…
Jetson-PI: Towards Onboard Real-Time Robot Control via Foresight-Aligned Asynchronous Inf…
NexForge: Scaling Agent Capabilities through Requirement-Driven Task Synthesis for LLMs
A Scaffolded GenAI Lab in Early Undergraduate CS: A Mixed-Methods, Multi-Course Evaluation
Beyond Success Rate: Cost-Aware Evaluation of Offensive and Defensive Security Agents
Memory-Driven Self-Disclosure and Relational Turning Points: A Longitudinal Multimodal St…
Digital Pantheon: Simulating and Auditing Coalition Formation with LLM Agents
ToolVerse: Unlocking Massive Environments and Long-Horizon Tasks for Agentic Reinforcemen…
Behavioral Controllability of Agentic Models for Information Extraction: From Fixed Workf…
A Formally Grounded ODRL Evaluator: Implementation and Comparison
CRAFT: Clustering Rubrics to Diagnose Weak LLM Capabilities and Generate Targeted Fine-Tu…
How Does Empowering Users with Greater System Control Affect News Filter Bubbles?
On the Structure of Address in Multi-Party Dialogue: From Discrete Labels to Continuous L…
Token-Level Off-Policy Learning for Faithful Generation Under Distribution Shift
DecoyFace: Beyond Obfuscation via Controllable and Imperceptible Identity Misdirection fo…
Panache: One-Pass Motif Discovery at Every Window Length
Bio-SFT: Asymmetric Cortical Guidance and Retinal Adaptation for Robust HDR Reconstruction