Real vs. Semi-Simulated: Rethinking Evaluation for Treatment Effect Estimation
Explorar
Noticias de IA
27413 elementos — filtrados, clasificados y sin duplicados
E = T*H/(O+B): A Dimensionless Control Parameter for Mixture-of-Experts Ecology
Quality-Conditioned Agreement in Automated Short Answer Scoring: Mid-Range Degradation an…
MoBayes: A Modular Bayesian Framework for Separating Reasoning from Language in Conversat…
EditCaption: Human-Refined SFT and HAE-DPO for Image Editing Instruction Synthesis
Inference Time Context Sparsity: Illusion or Opportunity?
Theoretical Analysis of Sparse Optimization with Reparameterization, Weight Decay, and Ad…
Identifying and Mitigating Systemic Measurement Bias in Production LLM Inference Benchmar…
How Well Do Models Follow Their Constitutions?
HiGraph: A Large-Scale Hierarchical Graph Dataset for Malware Analysis
Prism: Spectral-Aware Block-Sparse Attention
Scalable Explainability-as-a-Service (XaaS) for Edge AI Systems
SPA-Cache: Singular Proxies for Adaptive Caching in Diffusion Language Models
Dynamics Reveals Structure: Challenging the Linear Propagation Assumption
Characterizing Linear Alignment Across Language Models
Reliable AI Needs to Externalize Implicit Knowledge: A Human-AI Collaboration Perspective
Probing the Preferences of a Language Model: Integrating Verbal and Behavioral Tests of A…
From Prompt Optimization to Multi-Dimensional Credibility Evaluation: Enhancing Trustwort…
Learning to Trust: Bayesian Adaptation to Varying Suggester Reliability in Sequential Dec…
SPARK: Search Personalization via Agent-Driven Retrieval and Knowledge-sharing
PathWise: Planning through World Model for Automated Heuristic Design via Self-Evolving L…
LGMT: Logic-Grounded Metamorphic Testing for Evaluating the Reasoning Reliability of LLMs
E3AD: An Emotion-Aware Vision-Language-Action Model for Human-Centric End-to-End Autonomo…
Why We Need World Models for AGI: Where LLMs Fail and How World Models May Outperform
BackWeak: Backdooring Knowledge Distillation Simply with Weak Triggers and Fine-tuning
LC-ERD: Mining Latent Logic for Self-Evolving Reasoning via Consistency-Regulated Reward …
Breaking the Chains of Probability: Neutrosophic Logic as a New Framework for Epistemic U…
SkillEvolBench: Benchmarking the Evolution from Episodic Experience to Procedural Skills
ChunkLLM: A Lightweight Pluggable Framework for Accelerating LLMs Inference
Go witheFlow: Real-time Emotion Driven Audio Effects Modulation