How we run inference today — NVIDIA GPUs on AWS, measured, not a private cluster
What runs today, what has been measured on NVIDIA GPUs, and what the public demos still use. Written as facts, because that is what they are.
Start with what actually runs. Today QUOKKA AI operates a public site, an internal CMS, and an API: Nuxt 4 on the front, Vue 3 in the CMS, ThinkPHP 8 behind it, MySQL 8.4 underneath, with automated Playwright smoke suites gating every change. Public MindLog and Pulse demos still call a hosted API. There is no 24/7 inference endpoint on our own GPUs, and no private cluster. A company page that claimed otherwise would be easy to check and wrong.
NVIDIA GPU inference has been measured. On AWS g4dn.xlarge in Sydney — NVIDIA Tesla T4, 16GB — with CUDA and vLLM 0.27.1, serving Qwen2.5-7B-Instruct-AWQ. First token about 34 ms; about 35 tokens per second on a single stream. TensorRT, Triton, NVIDIA NIM, and NeMo are not in use. The instance is on-demand, not a standing inference cluster. Larger G5/G6 instances when Sydney capacity allows.
Around the model, the target product runtime is still ordinary AWS: ECS for containerised product services, S3 for traces and assets, RDS for relational state, CloudFront at the edge, CloudWatch for operational visibility. Session and content state sit on MySQL now; vector memory, caches, and versioned traces land with the runtime rather than before it.
What will not change is the shape of the company. QUOKKA AI is not a consulting shop, a cloud reseller, or a GPU-hours broker. Usage, once products ship, will be product usage — MindLog sessions, Pulse API requests — billed under a published plan. The in-house work is orchestration: how work is split, how agents hand off, how memory lasts across a run, and how scoring and guardrails decide what is allowed to reach a person. The GPUs are a means to that, not the pitch.