Bridging Evolutionary Algorithms and Reinforcement Learning: A Comprehensive Survey on Hy…
Explorar
Noticias de IA
37834 elementos — filtrados, clasificados y sin duplicados
Fusion Embedding for Pose-Guided Person Image Synthesis with Diffusion Model
Filtered Posterior Mean Collections: A Unified Framework for Analytical Models of Diffusi…
Pragmatic Reasoning improves LLM Code Generation
PhySense: Sensor Placement Optimization for Accurate Physics Sensing
Reward-free Alignment for Conflicting Objectives
MultiPhishGuard: An Explainable and Adaptive Multi-Agent LLM System for Phishing Email De…
ClaimDiff-RL: Fine-Grained Caption Reinforcement Learning through Visual Claim Comparison
SoK: A Comprehensive Security Analysis of Jailbreak Resilience in GPT and DeepSeek Models
FLoRIST: Singular Value Thresholding for Efficient and Accurate Federated Fine-Tuning of …
Agent Primitives: Reusable Latent Building Blocks for Multi-Agent Systems
Understanding Data Temporality Impact on Large Language Models Pre-training
Adaptive Human-AI Coordination via Hierarchical Action Disentanglement
Beyond Final Answers: Auditing Trajectory-Level Hallucinations in Multi-Agent Industrial …
TimeGuard: Channel-wise Pool Training for Backdoor Defense in Time Series Forecasting
A Sober Look at Agentic Misalignment in Automated Workflows
MAPLE: Multi-State Aggregated Policy Evaluation for AlphaZero in Imperfect-Information Ga…
Look-Closer-Then-Diagnose: Confidence-Aware Ultrasound VQA via Active Zooming
The Devil is in the Condition Numbers: Why is GLU Better than non-GLU Structure?
BEAR: Towards Beam-Search-Aware Optimization for Recommendation with Large Language Models
CounterFlow: A Two-Phase Inference-Time Sampling for Counterfactual Video Foley Generation
A Simplex Witness Certificate for Constant Collapse in Variational Autoencoders
AME-TS: Anchored Mixture-of-Experts for Time Series Forecasting
RecGOAT: Graph Optimal Adaptive Transport for LLM-Enhanced Multimodal Recommendation with…
BacktestBench: Benchmarking Large Language Models for Automated Quantitative Strategy Bac…
Learning Selective Merge Policies for Deadline-Constrained Coded Caching via Deep Reinfor…
HEPA: A Self-Supervised Horizon-Conditioned Event Predictive Architecture for Time Series
Intelligent Truck Matching in Full Truckload Shipments using Ping2Hex approach
Internalizing Outcome Supervision into Process Supervision: A New Paradigm for Reinforcem…
Sparse Tokens Suffice: Jailbreaking Audio Language Models via Token-Aware Gradient Optimi…