Pendra
Run open-source AI models on your own GPUs or UK-hosted hardware, with one OpenAI-compatible API and zero data retention.
| What is it | Run open-source AI models on your own GPUs or UK-hosted hardware, with one OpenAI-compatible API and zero data retention. |
|---|---|
| Pricing | Freemium — from £99/mo |
| Free tier | Yes |
| Platform | Web Application |
| API | Yes |
| Best for | Clinical document processing (summarise records, extract from discharge notes), Privileged document review in legal firms |
| Domain registered | 2026 |
Data updated Aug. 5, 2026
What does Pendra do?
Pendra is a managed inference platform built for teams in regulated industries that need to use AI without handing sensitive data to a third party. Instead of sending prompts to a public cloud, you install a small worker on your own GPUs (or use Pendra's UK-hosted ones), pull open-weight models like Qwen, Llama, or DeepSeek onto it, and call them through a single API. The whole setup is designed so your data never leaves your control — no retention, no foreign jurisdiction, no exposure.
Key features
What makes it stand outWho is Pendra for?
Who benefits most from this toolPricing
Free tier available — start without a credit cardFree
- Personal use only
- 1 self-hosted Pendra Worker
- OpenAI-compatible API endpoint
- Zero data retention
- Sovereign jurisdiction
- Python and Node.js SDKs
- Community documentation
Pro
Everything in Free, plus:
- Commercial use rights
- Up to 5 self-hosted Pendra Workers
- Enhanced usage analytics
- Worker monitoring & webhook alerts
- Standard DPA included
- Priority support (email)
- Request logging (optional)
- Private inference (end-to-end encryption)
Enterprise
Everything in Pro, plus:
- Pendra-managed GPU infrastructure
- Unlimited self-hosted Workers
- Automatic data masking and redaction
- Granular audit logging
- Custom DPA and DPIA support
- Dedicated account manager
- SLA with uptime guarantee
- SSO and role-based access control
- Onboarding and integration support
Trust & presence
Gallery
Click any image to enlargeAlternatives in AI inference
Cloud platform that rents NVIDIA H100/B200/B300 GPUs for training and running AI models at scale
Serverless GPU platform for deploying machine learning models in minutes, with auto-scaling and pay-per-use pricing.
Self-hosted AI inference engine for search and document processing — deploy models on your own cloud infrastructure.
Optimize open-source AI models for production — benchmark engines, tune latency, and deploy on any GPU.
Cloud GPU platform for AI developers — deploy, train, and scale AI models with on-demand infrastructure
High-throughput LLM inference engine for fast, memory-efficient AI model serving.
Cerebras’ third-generation wafer-scale engine (WSE-3) is the fastest AI processor on Earth. It surpasses all other processors in AI-optimized cores, memory speed, and on-chip fabric bandwidth.
Rent dedicated GPU servers and VPS for AI, rendering, and LLM hosting, starting at $85/month.