Multi-agent products — in design
Surfaces for life replay, scoring, and generative worlds. MindLog and Pulse are the first two, both in early build.
Technology
The public site, the CMS, and the API are running now: Nuxt 4, Vue 3, ThinkPHP 8, MySQL 8.4. NVIDIA GPU inference has been measured on AWS. Public MindLog and Pulse demos still call a hosted API. This page keeps those facts apart.
View the architecture boardAI Stack
Surfaces for life replay, scoring, and generative worlds. MindLog and Pulse are the first two, both in early build.
Roles, handoffs, memory, and protocols designed in-house — not left to a single chat turn. This is the current build focus.
Measured on AWS g4dn.xlarge in Sydney (NVIDIA Tesla T4, 16GB) with CUDA and vLLM. Model: Qwen2.5-7B-Instruct-AWQ. On-demand, not a 24/7 inference cluster. No private training cluster, now or planned. Larger G5/G6 instances when Sydney capacity allows.
Serving path measured for the MVP: CUDA with vLLM 0.27.1. TensorRT and Triton are not in use. NVIDIA NIM and NeMo are not in use.
Embeddings and durable state to keep generation anchored in what just happened. Session and content state today sit on MySQL.
Planning, tool use, handoffs, and recovery for long-running work. This is what is being written right now.
Automated Playwright smoke suites already gate the site, CMS, and API on every change. Product scoring, replay, and safety suites land with the runtime.
On T4 with vLLM: first token ~34 ms; ~35 tokens/s on a single stream; ~139 tokens/s at batch 4. Public demos still use a hosted API.
System Architecture
Web today. Mini-app, mobile, and H5 surfaces follow as products ship.
Planned: routing, step quotas, auth, and streaming at the edge.
Runtime, memory, agent orchestration, and scoring — designed as separate services, in build now.
Session and content state run on MySQL 8.4 today. Vector memory, caches, and versioned traces arrive with the runtime.
NVIDIA inference has been measured on GPU-backed EC2 (g4dn.xlarge, Tesla T4). ECS, S3, RDS, and CloudFront remain the target product runtime. Today the site, CMS, and API run from a single origin behind a CDN. Public demos still call a hosted API.
Measured on NVIDIA T4
Measured on demand on AWS. Not production traffic.
| Reading | Value |
|---|---|
| First token | ~34 ms |
| Single stream | ~35 tokens/s |
| Batch of 4 | ~139 tokens/s |
Core Capabilities