ScientistOne: Towards Human-Level Autonomous Research via Chain-of-Evidence
Browse
AI News
30934 items — filtered, classified, deduplicated
JobBench: Aligning Agent Work With Human Will
Post-training makes large language models less human-like
Governed Metaprogramming for Intelligent Systems: Reclassifying Eval as a Governed Effect
The Strongest Teacher Is Not Always the Best Teacher: Student-Centric Answer Selection
Persistent AI Agents in Academic Research: A Single-Investigator Implementation Case Study
Innovation: An Almost Characterization of Hallucination
Symmetry-Compatible Principle for Optimizer Design: Embeddings, LM Heads, SwiGLU MLPs, an…
Traceable Knowledge Graph Reasoning Enables LLM-Assisted Decision Support for Industrial …
ICCU: In-Context Continual Unlearning via Pattern-Induced Refusal Rules
VIDA: A dataset for Visually Dependent Ambiguity in Multimodal Machine Translation
Generative artificial intelligence and the marginalization of minoritized knowledges in h…
Tracing the Dynamics of Refusal: Exploiting Latent Refusal Trajectories for Robust Jailbr…
CyberEvolver: Structured Self-Evolution for Cybersecurity Agents On the Fly
Measuring Prediction Uncertainty in Neural Cellular Automata
Beyond Trajectory-Level Attribution: Graph-Based Credit Assignment for Agentic Reinforcem…
Rethinking the Trust Region in LLM Reinforcement Learning
RulePlanner: All-in-One Reinforcement Learner for Unifying Design Rules in 3D Floorplanni…
DynFrame: Adaptive Reasoning-Driven Multimodal Framework with Dynamic Frame Augmentation …
More Expressive Feedforward Layers: Part I. Token-Adaptive Mixing of Activations
Spend Your Rollouts Where It Counts: Rollout Allocation for Group-Based RL Post-Training
Examining the Challenges of Intellectual Property in AI-Generated Productions
SWAP: Towards Copyright Auditing of Soft Prompts via Sequential Watermarking
Monte Carlo Permutation Search
When LLMs Benchmark Themselves: Deconstructing Self-Bias in Automated Evaluation
A Physics-Informed Hierarchical Neural Network for Microwave Scattering Analysis of 3D PE…
Adaptive Multi-prompt Contrastive Network for Few-shot Out-of-distribution Detection
Personalizing Embodied Multimodal Large Language Model Agents over Long-term User Interac…
Can LLMs Introspect? A Reality Check
Bridging Control with Neural Network Verifier alpha-beta-CROWN: A Tutorial