What runs today, and what has been measured on NVIDIA GPUs.

The public site, the CMS, and the API are running now: Nuxt 4, Vue 3, ThinkPHP 8, MySQL 8.4. NVIDIA GPU inference has been measured on AWS. Public MindLog and Pulse demos still call a hosted API. This page keeps those facts apart.

View the architecture board

Running today, and the layers the first MVP still needs.

01In design

Multi-agent products — in design

Surfaces for life replay, scoring, and generative worlds. MindLog and Pulse are the first two, both in early build.

02In build

Orchestration architecture

Roles, handoffs, memory, and protocols designed in-house — not left to a single chat turn. This is the current build focus.

03Measured

NVIDIA GPU inference — measured

Measured on AWS g4dn.xlarge in Sydney (NVIDIA Tesla T4, 16GB) with CUDA and vLLM. Model: Qwen2.5-7B-Instruct-AWQ. On-demand, not a 24/7 inference cluster. No private training cluster, now or planned. Larger G5/G6 instances when Sydney capacity allows.

04In use

CUDA · vLLM — in use

Serving path measured for the MVP: CUDA with vLLM 0.27.1. TensorRT and Triton are not in use. NVIDIA NIM and NeMo are not in use.

  • TensorRT— Not in use
  • Triton— Not in use
  • NVIDIA NIM— Not in use
  • NeMo— Not in use
05Planned

Memory and retrieval — planned

Embeddings and durable state to keep generation anchored in what just happened. Session and content state today sit on MySQL.

06In build

Agent orchestration — in build

Planning, tool use, handoffs, and recovery for long-running work. This is what is being written right now.

07In use

Evaluation harness

Automated Playwright smoke suites already gate the site, CMS, and API on every change. Product scoring, replay, and safety suites land with the runtime.

08Measured

Real-time experience — measured on T4

On T4 with vLLM: first token ~34 ms; ~35 tokens/s on a single stream; ~139 tokens/s at batch 4. Public demos still use a hosted API.

Architecture: what runs today, what is targeted

  • Running today
  • Measured, not in production
  • In build
  • Planned
  1. Clients
    • Web — Running today
    • Mini-Apps — Planned
    • Mobile — Planned
    • H5 / Webview — Planned
  2. Orchestration gateway
    • Agent routing — Planned
    • Auth — Planned
    • Step quotas — Planned
    • Streaming transport — Planned
  3. Product services
    • Runtime — In build
    • Memory — In build
    • Agent runtime — In build
    • Scoring — In build
  4. Data Layer
    • MySQL 8.4 · today — Running today
    • KV cache · planned — Planned
    • Vector store · planned — Planned
    • Trace store · planned — Planned
  5. Runtime — target on AWS
    • Amazon EC2 GPU · T4 measured — Measured, not in production
    • Amazon ECS · planned — Planned
    • S3 · RDS · CloudFront · planned — Planned
    • CloudWatch / traces · planned — Planned
  1. Client layer

    Web today. Mini-app, mobile, and H5 surfaces follow as products ship.

  2. Orchestration gateway

    Planned: routing, step quotas, auth, and streaming at the edge.

  3. Product services

    Runtime, memory, agent orchestration, and scoring — designed as separate services, in build now.

  4. Data layer

    Session and content state run on MySQL 8.4 today. Vector memory, caches, and versioned traces arrive with the runtime.

  5. Target runtime (AWS)

    NVIDIA inference has been measured on GPU-backed EC2 (g4dn.xlarge, Tesla T4). ECS, S3, RDS, and CloudFront remain the target product runtime. Today the site, CMS, and API run from a single origin behind a CDN. Public demos still call a hosted API.

Three readings, and nothing we have not measured.

Measured on demand on AWS. Not production traffic.

GPU
NVIDIA Tesla T4 · 16 GB
Setup
AWS g4dn.xlarge · Sydney
Serving
CUDA · vLLM 0.27.1
Model
Qwen2.5-7B-Instruct-AWQ
T4 inference readings — Measured on demand on AWS. Not production traffic.
ReadingValue
First token~34 ms
Single stream~35 tokens/s
Batch of 4~139 tokens/s

Where the work is, and what is still ahead

In-house orchestration
How work is split, how agents collaborate, and how state is remembered — designed and built in-house. This is the current build.
GPU cost discipline
Batching, KV-cache reuse, and vLLM keep T4 cost inside a product budget. We do not sell or resell GPU hours.
Real-time latency
Measured on T4 with vLLM: first token ~34 ms. Public product demos still stream from a hosted API.
Evaluation discipline
Automated smoke suites already gate the public site, CMS, and API. Product scoring, replay, and safety gates arrive with the runtime.
Product safety
Guardrails, review paths, and constraints designed into the runtime as it is built.
Trace flywheel — planned
Capture and curate production traces once products serve real sessions, then feed them back into the next orchestration pass.