Service-Induced Congestion in Memory-Constrained LLM Serving
Explorar
Noticias de IA
676 elementos — filtrados, clasificados y sin duplicados
AI Agent Failure Detection and Root Cause Analysis with Strands Evals
Nvidias RTX Spark is big news, but its not for everyone
My Homelab AI Dev Platform
Build context-rich research agents with Deep Agents and Bedrock AgentCore
A Benchmark and Framework for Evaluating Next Action Predictions in Spreadsheets
MASLab: A Unified and Comprehensive Codebase for LLM-based Multi-Agent Systems
STREAM: Multi-Tier LLM Inference Middleware with Dual-Channel HPC Token Streaming
AI Coding at Home Without Going Broke
CODA-BENCH: Can Code Agents Handle Data-Intensive Tasks?
olmo-eval: An evaluation workbench for the model development loop
Extract Data with On-demand and Batch Pipelines Dynamically
Evaluate AI agents systematically with Agent-EvalKit
Show HN: Fata – Spaced repetition to fight skill rot from AI coding
Layer-Isolated Evaluation: Gating the Deterministic Scaffold of a Production LLM Agent wi…
Profiling in PyTorch (Part 2): From nn.Linear to a Fused MLP
Stop hand-tuning kernels: How Neuron Agentic Development accelerates AWS Trainium optimiz…
Apache Burr: Build reliable AI agents and applications
Piper: A Programmable Distributed Training System
Towards Autonomous Accelerator Design: FPGA Accelerator Generation with SECDA
TinyTroupe: An LLM-powered Multiagent Persona Simulation Toolkit
torch-sla: Differentiable Sparse Linear Algebra with Adjoint Solvers and Sparse Tensor Pa…
Scale Robot Reinforcement Learning with NVIDIA Isaac Lab on Amazon SageMaker AI
Hands-free first notice of loss: Using Strands Agents and Amazon Bedrock AgentCore Browse…
Build an agentic incident triage assistant with Amazon Quick and New Relic
Build a Basic AI Agent from Scratch: Long Task Planning
How engineers at Nextdoor use Codex to build without limits
How an Agent Built a 3D Paris Gallery by Chaining Two Hugging Face Spaces
Context Rot in AI-Assisted Software Development: Repurposing Documentation Consistency fo…
vla.cpp: A Unified Inference Runtime for Vision-Language-Action Models