GrepSeek: Training Search Agents for Direct Corpus Interaction
Explorar
Noticias de IA
27413 elementos — filtrados, clasificados y sin duplicados
OpenClawBench: Benchmarking Process-side Anomalies in Real-world Agent Execution Trajecto…
Surfacing Isolated Learners with Outcome-Independent Mediation of Feedback between Teache…
AtomWorld: A Benchmark for Evaluating Spatial Reasoning in Large Language Models on Cryst…
Benchmarking LLM-Assisted Blue Teaming via Standardized Threat Hunting
CoHyDE: Iterative Co-Training of LLM Rewriter & Dense Encoder for Tool Retrieval
Better Later Than Sooner: Neuro-Symbolic Knowledge Graph Construction via Ontology-ground…
Scaling Monosemanticity: Extracting Interpretable Features from Claude 3 Sonnet
GroundAct: Can LLM Agents Ground Actions in Environmental States?
MENTOR: Efficient Multimodal-Conditioned Tuning for Autoregressive Vision Generation Mode…
HD-Prot: A Protein Language Model for Joint Sequence-Structure Modeling with Continuous S…
Position: Text Embeddings Should Capture Implicit Semantics, Not Just Surface Meaning
EPiC: Efficient Video Camera Control Learning with Precise Anchor-Video Guidance
GPIC: A Giant Permissive Image Corpus for Visual Generation
Steering Language Models Before They Speak: Logit-Level Interventions
The Confidence Shortcut: A Reasoning Failure Mode of Masked Diffusion Models
Dataset-Driven Channel Masks in Transformers for Multivariate Time Series
Aligned but Fragile: Enhancing LLM Safety Robustness via Zeroth-Order Optimization
Towards Human-Like Interactive Speech Recognition With Agentic Correction and Semantic Ev…
DeepSurvey: Enhancing Analytical Depth and Citation Reliability in Automated Survey Gener…
A Survey on Recent Advances in Conversational Data Generation
A Language-Guided Bayesian Optimization for Efficient LoRA Hyperparameter Search
ParaTool: Shifting Tool Representations from Context to Parameters
GPS-Enhanced Tourist Mobility Modeling with Seasonal Spatial Priors and LLM-Based Activit…
Beyond Normalization: Rethinking the Partition Function as a Difficulty Scheduler for RLVR
Reasoning Theater: Disentangling Model Beliefs from Chain-of-Thought
MPDocBench-Parse: Benchmarking Practical Multi-page Document Parsing
When Should a Robot Think? Resource-Aware Reasoning via Reinforcement Learning for Embodi…
Maximizing Mutual Information Between Prompt and Response Improves LLM Performance With N…
Revisiting the Effectiveness of LLM Pruning for Test-Time Scaling