PetroBench: A Benchmark for Large Language Models in Petroleum Engineering
Explorar
Noticias de IA
27413 elementos — filtrados, clasificados y sin duplicados
MemCog: From Memory-as-Tool to Memory-as-Cognition in Conversational Agents
MolLingo: Molecule-Native Representations for LLM-Powered Scientific Agents
FundaPod: A Multi-Persona Agent Pod Platform with Knowledge Graph Memory for AI-Assisted …
Towards Faithful Agentic XAI: A Verification Method and an Open-World Benchmark for Bette…
Dr-CiK: A Testbed for Foresight-Driven Agents
DiagramRAG: A Lightweight Framework to Retrieve Scientific Diagram for Figure Generation
The Shape of Overthinking: Backtracking Bursts in Long Reasoning Traces
Transferable Reinforcement Learning via Probabilistic Latent Embeddings and Dynamic Polic…
Sense Representations Are Inducible Interfaces
Identifying and Understanding Human Values in Text: A Tailorable LLM-based Architecture
Behavioural Analysis of Alignment Faking
Why LLMs Fail at Causal Discovery and How Interventional Agents Escape
RULER: Representation-Level Verification of Machine Unlearning
Cyberbullying Governance on Social Media: A Unified Framework from Content Identification…
A Policy-Driven Runtime Layer for Agentic LLM Serving
Got a Secret? LLM Agents Can't Keep It: Evaluating Privacy in Multi-Agent Systems
Revealing Algorithmic Deductive Circuits for Logical Reasoning
When Context Flips, Safety Breaks: Diagnosing Brittle Safety in Aligned Language Models
AIBuildAI-2: A Knowledge-Enhanced Agent for Automatically Building AI Models
PortBench: A Correlation-Aware, Full-Pipeline Benchmark for LLM-Driven Portfolio Manageme…
SKILLC: Learning Autonomous Skill Internalization in LLM Agents via Contrastive Credit As…
Show, Don't TELL: Explainable AI-Generated Text Detection
Harness-Bench: Measuring Harness Effects across Models in Realistic Agent Workflows
Do Agents Think Deeper? A Mechanistic Investigation of Layer-Wise Dynamics in Sequential …
Geometry of Human Perceptual Domains Emerges Transiently in LLM Representations
AsyncTool: Evaluating the Asynchronous Function Calling Capability under Multi-Task Scena…
Confidence-Orchestrated Self-Evolution against Uncertain LLM Feedback
Clark Hash: Stateless Sparse Johnson-Lindenstrauss Quantization for Neural Embeddings
Relevant Is Not Warranted: Evidence-Force Calibration for Cited RAG