kapyn
RadarInference

Cactus

Inference · Tool

Visit Website

About Cactus

A hybrid edge-cloud inference engine with OpenAI-compatible APIs, custom ARM kernels, and 1-4-bit quantization, so models can run locally on mobile hardware. Native bindings cover Swift, Kotlin, Flutter, React Native, Python, and Rust. Fills the on-device slot in an otherwise cloud-only inference field.

Run text, speech, and vision models on-device on phones and wearables.

CategoryInference
TypeTool
Websitegithub.com
See Cactus alternatives

Similar tools

OpenRouterOne API for hundreds of models, with fallback routing.GroqRun open models at extreme speed on custom hardware.Together AIHost and fine-tune open models via a fast API.ReplicateRun and deploy any model with one API call.Hugging FaceThe hub for open models, datasets, and demos.Fireworks AIFast, production inference for open models.

Find more on the Radar

Browse every tool, model, and MCP server in one place.

Open the Radar