AI inference — AI tools

262 tools in this category

Infrastructure for running models in production: serve checkpoints behind an API, rent GPUs by the second, batch large jobs, and watch latency and cost per request. Sits after the training stage — this is where a model becomes something an app can call.

Works with
262 tools
NVIDIA Verified Developer Tools

AI computing platform providing GPUs, software, and infrastructure for training and deploying AI models across industries

Top 1k site
Nebius Verified AI inference

AI cloud platform providing scalable NVIDIA GPU clusters for training and inference, with managed Kubernetes and Slurm orchestration.

Groq Verified AI inference

High-speed, low-cost AI inference API for running large language models with minimal latency.

zapier · make+2 Top 100k site
Vast ai Verified AI inference

Rent high-performance GPUs on demand for AI, machine learning, and graphics rendering at significantly lower costs.

Top 100k site
Lambda Verified AI inference

Cloud platform that rents NVIDIA H100/B200/B300 GPUs for training and running AI models at scale

Roboflow Verified AI inference

Organize images, convert annotation formats, preprocess, augment, share, and ship more. We eliminate the boilerplate code every computer vision team has to write.

n8n Top 100k site
RunPod Verified AI inference

Cloud GPU platform for AI developers — deploy, train, and scale AI models with on-demand infrastructure

Top 100k site
Dynatrace Verified Developer Tools

AI-powered observability platform that monitors apps, infrastructure, and security in real-time

n8n · workato Top 1k site
Baseten Verified AI inference

AI infrastructure platform — deploy, serve, and scale machine learning models in production with optimized inference

n8n Top 100k site
Cursor Verified Developer Tools

AI-powered code editor that helps you write, understand, and debug code faster with intelligent assistance.

n8n · ifttt Top 10k site
Sambanova Verified AI inference

Enterprise AI platform providing high-performance inference for large language models and agentic AI workflows.

TensorFlow Verified Developer Tools

End-to-end open source platform for building and deploying machine learning models across diverse environments

Top 100k site
Modal Verified Developer Tools

Run or deploy machine learning models, massively parallel compute jobs, task queues, web apps, and much more, without your own infrastructure.

Ultralytics Verified No-Code&Low-Code

No-code platform for training and deploying custom computer vision models — upload images, select a model, and deploy to any device.

Top 100k site
Zenlayer Verified AI inference

Distributed cloud platform for deploying and scaling AI inference and compute globally

n8n
Cerebras Verified AI inference

Cerebras’ third-generation wafer-scale engine (WSE-3) is the fastest AI processor on Earth. It surpasses all other processors in AI-optimized cores, memory speed, and on-chip fabric bandwidth.

Top 100k site
Llama Verified LLM

Open-source multimodal AI models for developers — build apps with text, image, and long-context capabilities

make Top 100k site
vLLM Verified AI inference

High-throughput LLM inference engine for fast, memory-efficient AI model serving.

Top 100k site
Crusoe Verified AI inference

Renewable-powered cloud infrastructure and managed inference service for running large AI models.

Ollama Verified LLM

Run and manage large language models locally on your machine for private, secure AI automation.

n8n · workato Top 100k site
Hailo AI Verified AI inference

Edge AI processors that enable high-performance deep learning applications on devices at ultra-low power consumption.

Cohere Verified LLM

Enterprise AI platform for building secure, customizable language models that run on your own infrastructure.

Top 100k site
Chutes Verified AI inference

Serverless AI compute platform for running open-source LLMs, image, video, and audio models at scale

n8n
fireworks.ai Verified AI inference

High-performance AI inference platform — deploy and scale open-source models like Llama and Gemma with a single API call.

More in Development