CloudTwin
Overview
CloudTwin creates a live virtual counterpart — a digital twin — for every Docker container running on a monitored host. Each twin continuously mirrors its container's real-time state, computes a health score, predicts failure risk, and can autonomously propose and execute self-healing actions through a human-in-the-loop approval workflow.
Most observability tools stop at Monitor → Alert. CloudTwin implements a full closed loop:
Monitor → Predict → Decide → Act → Confirm
| Problem | CloudTwin's Solution |
|---|---|
| Reactive alerting | Rules engine evaluates sustained thresholds, creates alerts before users notice |
| Manual remediation | Decision engine proposes restarts; agent executes on human approval |
| No trend awareness | What-if simulator projects metrics at increased load using historical slope |
| Opaque infrastructure | Health score (0–100) synthesises CPU + memory into a single readable signal |
| No accountability | Every alert, action, and ack recorded with full lifecycle timestamps |
| Tight coupling | Agent-pull pattern — backend never needs network access to target host |
| Multi-user chaos | Workspace isolation — every user sees only their own containers |
Screenshots
System Architecture

Dashboard Overview

What-if Simulation

Alerts Feed

Actions Panel

Login

Settings

Live Docker Containers (Agent monitoring)

Data Model

Architecture
The system is structured across three independent planes:
Agent plane — A compiled Go binary runs on the monitored host. It polls the local Docker Engine API via Unix socket every 10 seconds, builds metric snapshots, and POSTs them directly to the backend over HTTPS using an API key. A second goroutine polls the backend for approved commands and executes them locally via Docker — the backend never connects to the host directly.
Backend plane — A Node.js/Express server receives snapshots via POST /api/snapshots, runs each through the twin engine (health score computation, Postgres upsert), evaluates workspace-scoped rules, creates alerts, and proposes restart actions. All data is isolated by workspace ID derived from the agent's API key.
Control plane — A human approves or rejects proposed actions on the dashboard. Approved actions push to a per-workspace Redis queue. The agent polls, claims, and executes them, then acks the result. Every transition from pending → approved → executing → completed (or failed) is permanently recorded.
Key Design Decisions
HTTP-only agent — no MQTT required
The agent publishes snapshots directly to POST /api/snapshots over HTTPS. Users need only 3 environment variables — no MQTT broker URL, no shared credentials, no broker setup. This is how Datadog, Grafana Agent, and New Relic all work: the agent calls home to your API, internal routing is your concern.
# All a user needs to run the agent
export CLOUDTWIN_BACKEND_URL=https://cloudtwin.onrender.com
export CLOUDTWIN_API_KEY=ct_xxx
export CLOUDTWIN_WORKSPACE_ID=xxx
./cloudtwin-agent
Agent-pull command execution
The backend never SSHes or connects directly to a monitored host. Approved commands push to a Redis list keyed commands:<workspaceId>:<twinId>. The agent polls /commands/claim, pops the command, runs it locally via Docker socket, and acks the result. This is the same pattern used by Ansible, Teleport, and AWS SSM.
Workspace isolation
Every API key is scoped to a workspace. Every database query filters by workspace_id derived from the API key. Two users' containers never mix, and one user cannot access another's data even with knowledge of a twin ID.
Health score model
score = 100 - (cpu_percent × 0.5) - (memory_percent × 0.3)
≥ 80 → healthy
≥ 50 → warning
< 50 → critical
What-if simulation
slope = (latest_value - oldest_value) / elapsed_seconds
projected = current × multiplier + slope × 300s
ADD_REPLICA projected CPU ≥ 90% or memory ≥ 88%
MONITOR_CLOSELY projected CPU ≥ 70% or memory ≥ 75%
SAFE otherwise
Decision engine cooldown
Before proposing a restart, the engine checks for any non-terminal action on the same twin within the last 10 minutes. A flapping container won't flood the actions queue with duplicate proposals.
Tech Stack
Agent
| Technology | Purpose |
|---|---|
| Go 1.24 | Compiled binary, two independent goroutines (metrics + commands) |
| Docker Engine API (Unix socket) | Container enumeration, stats streaming, docker restart |
| Net/HTTP (stdlib) | Snapshot publishing + command polling over HTTPS |
Backend
| Technology | Version | Purpose |
|---|---|---|
| Node.js + Express | 4.x | REST API, snapshot ingestion, request routing |
bcryptjs | 2.4 | Password hashing |
jsonwebtoken | 9.0 | JWT auth for dashboard sessions |
express-validator | 7.2 | Input validation on all auth routes |
express-rate-limit | 7.4 | 300 req/min global, 20 req/15min on auth |
pg (node-postgres) | 8.13 | PostgreSQL connection pool with SSL |
redis | 4.7 | Per-workspace per-twin command queues |
Frontend
| Technology | Version | Purpose |
|---|---|---|
| Next.js | 16.2 | React framework, App Router |
| React | 19.2 | UI rendering |
| TypeScript | 5 | Type-safe API contracts |
| Tailwind CSS | 4 | Monochrome dark theme |
| Recharts | 3.9 | Live CPU sparklines per twin card |
| Lucide React | 0.383 | Icon system |
| Framer Motion | — | MacBook scroll animation on landing page |
Infrastructure (Production)
| Service | Provider | Purpose |
|---|---|---|
| Backend | Render (free) | Node.js web service |
| PostgreSQL | Render (free) | Primary database |
| Redis | Render (free) | Command queue |
| Frontend | Vercel (free) | Next.js hosting |
Data Model
users
├── id UUID PK
├── email TEXT UNIQUE
├── password_hash TEXT bcrypt, never stored raw
└── name TEXT
workspaces
├── id UUID PK
└── owner_id UUID FK → users
api_keys
├── id UUID PK
├── workspace_id UUID FK → workspaces
├── key_hash TEXT UNIQUE SHA-256, never stored raw
└── last_used_at TIMESTAMPTZ
twins
├── twin_id TEXT
├── workspace_id UUID FK → workspaces
├── name TEXT
├── status TEXT healthy | warning | critical | unknown
├── health_score INTEGER 0–100
├── state JSONB latest raw snapshot
└── PRIMARY KEY (twin_id, workspace_id)
metric_snapshots
├── twin_id TEXT
├── workspace_id UUID
├── captured_at TIMESTAMPTZ
├── cpu_percent DOUBLE
├── memory_percent DOUBLE
├── memory_mb DOUBLE
└── payload JSONB
rules
├── workspace_id UUID FK → workspaces
├── metric TEXT cpu | memory
├── operator TEXT > | >= | < | <=
├── threshold DOUBLE
└── severity TEXT warning | critical
alerts
├── workspace_id UUID
├── twin_id TEXT
├── rule_id INTEGER FK → rules
└── metric, current_value, threshold, severity, message
actions
├── workspace_id UUID
├── twin_id TEXT
├── action_type TEXT restart
├── status TEXT pending → approved → executing → completed | failed | rejected
└── approved_at, executed_at, completed_at TIMESTAMPTZ
Closed-Loop Flow
1. Agent POSTs snapshot POST /api/snapshots { containerId, cpuPercent: 95, ... }
│ X-API-Key: ct_xxx
▼
2. Backend derives workspace API key → workspace_id
│
▼
3. Twin engine upsertTwin → health_score = 2, status = critical
│
▼
4. Rules engine cpu > 90 matched → createAlert
│
▼
5. Decision engine cooldown clear → INSERT actions (status: pending)
│
▼
6. Dashboard ActionsPanel human sees RESTART proposed for <container>
│
┌────┴────┐
Approve Reject
│
▼
7. POST /actions/:id/approve RPUSH commands:<workspaceId>:<twinId>
│
▼
8. Agent command loop GET /commands/claim → pops command → status: executing
│
▼
9. docker restart Docker Engine API: POST /containers/:id/restart
│
▼
10. POST /commands/:id/ack { success: true } → status: completed
API Reference
Auth
| Method | Endpoint | Auth | Description |
|---|---|---|---|
POST | /auth/signup | — | Create account — returns token + agentApiKey (once) |
POST | /auth/login | — | Sign in — returns JWT |
GET | /auth/me | JWT | Current user + workspace ID |
POST | /auth/api-key/rotate | JWT | Rotate agent API key |
Twins
| Method | Endpoint | Auth | Description |
|---|---|---|---|
GET | /twins | JWT | All twins in workspace |
GET | /twins/:id | JWT | Single twin detail |
GET | /twins/:id/history | JWT | Metric history (last 100) |
Alerts & Actions
| Method | Endpoint | Auth | Description |
|---|---|---|---|
GET | /alerts | JWT | All alerts, most recent first |
GET | /actions | JWT | All actions in workspace |
POST | /actions/:id/approve | JWT | Approve — pushes to Redis queue |
POST | /actions/:id/reject | JWT | Reject — terminal state |
Agent (API key auth)
| Method | Endpoint | Auth | Description |
|---|---|---|---|
POST | /api/snapshots | API Key | Ingest container metrics snapshot |
GET | /commands/claim | API Key | Pop next approved command |
POST | /commands/:id/ack | API Key | Report execution result |
Simulation & Settings
| Method | Endpoint | Auth | Description |
|---|---|---|---|
GET | /simulate/:twinId?loadMultiplier=2 | JWT | Project metrics at load multiple |
GET | /settings/rules | JWT | List workspace alert rules |
PATCH | /settings/rules/:id | JWT | Toggle rule enabled/disabled |
Running Locally
Prerequisites
- Docker + Docker Compose
- Go 1.21+
- Node.js 18+
1. Start infrastructure
cd infra
# Create MQTT password file (Mosquitto, for local infra only)
docker run --rm -v $(pwd):/output eclipse-mosquitto:2 \
mosquitto_passwd -c -b /output/mosquitto.passwd cloudtwin cloudtwin
docker compose up -d
# Starts: PostgreSQL (5432), Redis (6379)
2. Backend
cd backend
cp .env.example .env # fill in your values
npm install
node src/db/init.js # creates tables + seeds 3 default rules
npm start # listens on :4000
3. Frontend
cd frontend
echo "NEXT_PUBLIC_API_BASE=http://localhost:4000" > .env.local
npm install
npm run dev # listens on :3000
4. Sign up + run the agent
- Open
http://localhost:3000→ Sign up - Copy your API key and workspace ID from the onboarding page
- Run the agent:
cd agent
export CLOUDTWIN_BACKEND_URL=http://localhost:4000
export CLOUDTWIN_API_KEY=ct_your_key_here
export CLOUDTWIN_WORKSPACE_ID=your_workspace_id
go run main.go
Your Docker containers appear in the dashboard within 10 seconds.
Environment variables
backend/.env
PGHOST=localhost
PGPORT=5432
PGUSER=cloudtwin
PGPASSWORD=cloudtwin
PGDATABASE=cloudtwin
REDIS_URL=redis://localhost:6379
JWT_SECRET=generate-with-node-crypto-randomBytes-64-hex
CORS_ORIGIN=http://localhost:3000
frontend/.env.local
NEXT_PUBLIC_API_BASE=http://localhost:4000
Production Deployment
| Service | Platform | Notes |
|---|---|---|
| Backend | Render Web Service | Auto-deploys from main |
| PostgreSQL | Render Postgres | Free tier — export before 90-day limit |
| Redis | Render Redis | Free tier — no persistence |
| Frontend | Vercel | Auto-deploys from main, always-on |
Deploy your own
- Fork this repo
- Create Render Postgres + Redis in Singapore region
- Create Render Web Service → root directory
backend→ runtime Node → add all env vars - Run
node src/db/init.jsfrom Render Shell to apply schema - Import repo to Vercel → root directory
frontend→ addNEXT_PUBLIC_API_BASE - Update
CORS_ORIGINin Render env to your Vercel URL
Current Status
| Feature | Status |
|---|---|
| Docker Engine API metric collection | ✅ |
| HTTP snapshot ingestion (no MQTT for users) | ✅ |
| JWT auth + API key auth | ✅ |
| Workspace isolation (multi-user) | ✅ |
| Twin engine with health score | ✅ |
| Rules-based alert engine | ✅ |
| Decision engine with cooldown | ✅ |
| Redis command queue (workspace-scoped) | ✅ |
| Agent-pull self-healing execution | ✅ |
| Full action lifecycle audit trail | ✅ |
| What-if load simulation | ✅ |
| Input validation + rate limiting | ✅ |
| Jest test suite (9 tests) | ✅ |
| GitHub Actions CI | ✅ |
| Render + Vercel deployment | ✅ |
| WebSocket live push | ❌ Polling every 5s |
| Kubernetes agent | ❌ Planned |
| ML anomaly detection | ❌ Planned |
Planned Extensions
- Kubernetes agent — watch Pod resources via K8s API,
kubectl rollout restartas the action type - ML anomaly detection — Isolation Forest per twin, replaces fixed thresholds with learned baselines
- WebSocket live push — Redis pub/sub → broadcaster, eliminates 5s polling lag
- Scale-replica action —
docker compose up --scalealongside restart
License
MIT — see LICENSE.


