CloudTwin landing page
Live

CloudTwin

A digital twin platform for containerized infrastructure.

Domain

cloud-twin.vercel.app

Stack

Digital TwinHuman-in-loop AutomationGoNext.jsExpressDockerRules Engine

CloudTwin

Real-Time Digital Twin Platform for Containerized Infrastructure

Go Node.js Next.js TypeScript PostgreSQL Redis Docker License

Live Demo · Report Bug


Overview

CloudTwin creates a live virtual counterpart — a digital twin — for every Docker container running on a monitored host. Each twin continuously mirrors its container's real-time state, computes a health score, predicts failure risk, and can autonomously propose and execute self-healing actions through a human-in-the-loop approval workflow.

Most observability tools stop at Monitor → Alert. CloudTwin implements a full closed loop:

Monitor → Predict → Decide → Act → Confirm
ProblemCloudTwin's Solution
Reactive alertingRules engine evaluates sustained thresholds, creates alerts before users notice
Manual remediationDecision engine proposes restarts; agent executes on human approval
No trend awarenessWhat-if simulator projects metrics at increased load using historical slope
Opaque infrastructureHealth score (0–100) synthesises CPU + memory into a single readable signal
No accountabilityEvery alert, action, and ack recorded with full lifecycle timestamps
Tight couplingAgent-pull pattern — backend never needs network access to target host
Multi-user chaosWorkspace isolation — every user sees only their own containers

Screenshots

System Architecture

CloudTwin Architecture

Dashboard Overview

Overview

What-if Simulation

Simulation

Alerts Feed

Alerts

Actions Panel

Actions

Login

Login

Settings

Settings

Live Docker Containers (Agent monitoring)

Docker Desktop

Data Model

Twin Entity Schema


Architecture

The system is structured across three independent planes:

Agent plane — A compiled Go binary runs on the monitored host. It polls the local Docker Engine API via Unix socket every 10 seconds, builds metric snapshots, and POSTs them directly to the backend over HTTPS using an API key. A second goroutine polls the backend for approved commands and executes them locally via Docker — the backend never connects to the host directly.

Backend plane — A Node.js/Express server receives snapshots via POST /api/snapshots, runs each through the twin engine (health score computation, Postgres upsert), evaluates workspace-scoped rules, creates alerts, and proposes restart actions. All data is isolated by workspace ID derived from the agent's API key.

Control plane — A human approves or rejects proposed actions on the dashboard. Approved actions push to a per-workspace Redis queue. The agent polls, claims, and executes them, then acks the result. Every transition from pending → approved → executing → completed (or failed) is permanently recorded.


Key Design Decisions

HTTP-only agent — no MQTT required

The agent publishes snapshots directly to POST /api/snapshots over HTTPS. Users need only 3 environment variables — no MQTT broker URL, no shared credentials, no broker setup. This is how Datadog, Grafana Agent, and New Relic all work: the agent calls home to your API, internal routing is your concern.

# All a user needs to run the agent
export CLOUDTWIN_BACKEND_URL=https://cloudtwin.onrender.com
export CLOUDTWIN_API_KEY=ct_xxx
export CLOUDTWIN_WORKSPACE_ID=xxx
./cloudtwin-agent

Agent-pull command execution

The backend never SSHes or connects directly to a monitored host. Approved commands push to a Redis list keyed commands:<workspaceId>:<twinId>. The agent polls /commands/claim, pops the command, runs it locally via Docker socket, and acks the result. This is the same pattern used by Ansible, Teleport, and AWS SSM.

Workspace isolation

Every API key is scoped to a workspace. Every database query filters by workspace_id derived from the API key. Two users' containers never mix, and one user cannot access another's data even with knowledge of a twin ID.

Health score model

score = 100 - (cpu_percent × 0.5) - (memory_percent × 0.3)

≥ 80  →  healthy
≥ 50  →  warning
< 50  →  critical

What-if simulation

slope     = (latest_value - oldest_value) / elapsed_seconds
projected = current × multiplier + slope × 300s

ADD_REPLICA      projected CPU ≥ 90% or memory ≥ 88%
MONITOR_CLOSELY  projected CPU ≥ 70% or memory ≥ 75%
SAFE             otherwise

Decision engine cooldown

Before proposing a restart, the engine checks for any non-terminal action on the same twin within the last 10 minutes. A flapping container won't flood the actions queue with duplicate proposals.


Tech Stack

Agent

TechnologyPurpose
Go 1.24Compiled binary, two independent goroutines (metrics + commands)
Docker Engine API (Unix socket)Container enumeration, stats streaming, docker restart
Net/HTTP (stdlib)Snapshot publishing + command polling over HTTPS

Backend

TechnologyVersionPurpose
Node.js + Express4.xREST API, snapshot ingestion, request routing
bcryptjs2.4Password hashing
jsonwebtoken9.0JWT auth for dashboard sessions
express-validator7.2Input validation on all auth routes
express-rate-limit7.4300 req/min global, 20 req/15min on auth
pg (node-postgres)8.13PostgreSQL connection pool with SSL
redis4.7Per-workspace per-twin command queues

Frontend

TechnologyVersionPurpose
Next.js16.2React framework, App Router
React19.2UI rendering
TypeScript5Type-safe API contracts
Tailwind CSS4Monochrome dark theme
Recharts3.9Live CPU sparklines per twin card
Lucide React0.383Icon system
Framer Motion—MacBook scroll animation on landing page

Infrastructure (Production)

ServiceProviderPurpose
BackendRender (free)Node.js web service
PostgreSQLRender (free)Primary database
RedisRender (free)Command queue
FrontendVercel (free)Next.js hosting

Data Model

users
├── id            UUID PK
├── email         TEXT UNIQUE
├── password_hash TEXT          bcrypt, never stored raw
└── name          TEXT

workspaces
├── id            UUID PK
└── owner_id      UUID FK → users

api_keys
├── id            UUID PK
├── workspace_id  UUID FK → workspaces
├── key_hash      TEXT UNIQUE   SHA-256, never stored raw
└── last_used_at  TIMESTAMPTZ

twins
├── twin_id       TEXT
├── workspace_id  UUID FK → workspaces
├── name          TEXT
├── status        TEXT          healthy | warning | critical | unknown
├── health_score  INTEGER       0–100
├── state         JSONB         latest raw snapshot
└── PRIMARY KEY (twin_id, workspace_id)

metric_snapshots
├── twin_id       TEXT
├── workspace_id  UUID
├── captured_at   TIMESTAMPTZ
├── cpu_percent   DOUBLE
├── memory_percent DOUBLE
├── memory_mb     DOUBLE
└── payload       JSONB

rules
├── workspace_id  UUID FK → workspaces
├── metric        TEXT          cpu | memory
├── operator      TEXT          > | >= | < | <=
├── threshold     DOUBLE
└── severity      TEXT          warning | critical

alerts
├── workspace_id  UUID
├── twin_id       TEXT
├── rule_id       INTEGER FK → rules
└── metric, current_value, threshold, severity, message

actions
├── workspace_id  UUID
├── twin_id       TEXT
├── action_type   TEXT          restart
├── status        TEXT          pending → approved → executing → completed | failed | rejected
└── approved_at, executed_at, completed_at TIMESTAMPTZ

Closed-Loop Flow

1. Agent POSTs snapshot         POST /api/snapshots  { containerId, cpuPercent: 95, ... }
        │  X-API-Key: ct_xxx
        ▼
2. Backend derives workspace    API key → workspace_id
        │
        ▼
3. Twin engine                  upsertTwin → health_score = 2, status = critical
        │
        ▼
4. Rules engine                 cpu > 90 matched → createAlert
        │
        ▼
5. Decision engine              cooldown clear → INSERT actions (status: pending)
        │
        ▼
6. Dashboard ActionsPanel       human sees RESTART proposed for <container>
        │
   ┌────┴────┐
Approve    Reject
   │
   ▼
7. POST /actions/:id/approve    RPUSH commands:<workspaceId>:<twinId>
        │
        ▼
8. Agent command loop           GET /commands/claim → pops command → status: executing
        │
        ▼
9. docker restart               Docker Engine API: POST /containers/:id/restart
        │
        ▼
10. POST /commands/:id/ack      { success: true } → status: completed

API Reference

Auth

MethodEndpointAuthDescription
POST/auth/signup—Create account — returns token + agentApiKey (once)
POST/auth/login—Sign in — returns JWT
GET/auth/meJWTCurrent user + workspace ID
POST/auth/api-key/rotateJWTRotate agent API key

Twins

MethodEndpointAuthDescription
GET/twinsJWTAll twins in workspace
GET/twins/:idJWTSingle twin detail
GET/twins/:id/historyJWTMetric history (last 100)

Alerts & Actions

MethodEndpointAuthDescription
GET/alertsJWTAll alerts, most recent first
GET/actionsJWTAll actions in workspace
POST/actions/:id/approveJWTApprove — pushes to Redis queue
POST/actions/:id/rejectJWTReject — terminal state

Agent (API key auth)

MethodEndpointAuthDescription
POST/api/snapshotsAPI KeyIngest container metrics snapshot
GET/commands/claimAPI KeyPop next approved command
POST/commands/:id/ackAPI KeyReport execution result

Simulation & Settings

MethodEndpointAuthDescription
GET/simulate/:twinId?loadMultiplier=2JWTProject metrics at load multiple
GET/settings/rulesJWTList workspace alert rules
PATCH/settings/rules/:idJWTToggle rule enabled/disabled

Running Locally

Prerequisites

  • Docker + Docker Compose
  • Go 1.21+
  • Node.js 18+

1. Start infrastructure

cd infra

# Create MQTT password file (Mosquitto, for local infra only)
docker run --rm -v $(pwd):/output eclipse-mosquitto:2 \
  mosquitto_passwd -c -b /output/mosquitto.passwd cloudtwin cloudtwin

docker compose up -d
# Starts: PostgreSQL (5432), Redis (6379)

2. Backend

cd backend
cp .env.example .env       # fill in your values
npm install
node src/db/init.js        # creates tables + seeds 3 default rules
npm start                  # listens on :4000

3. Frontend

cd frontend
echo "NEXT_PUBLIC_API_BASE=http://localhost:4000" > .env.local
npm install
npm run dev                # listens on :3000

4. Sign up + run the agent

  1. Open http://localhost:3000 → Sign up
  2. Copy your API key and workspace ID from the onboarding page
  3. Run the agent:
cd agent
export CLOUDTWIN_BACKEND_URL=http://localhost:4000
export CLOUDTWIN_API_KEY=ct_your_key_here
export CLOUDTWIN_WORKSPACE_ID=your_workspace_id
go run main.go

Your Docker containers appear in the dashboard within 10 seconds.

Environment variables

backend/.env

PGHOST=localhost
PGPORT=5432
PGUSER=cloudtwin
PGPASSWORD=cloudtwin
PGDATABASE=cloudtwin
REDIS_URL=redis://localhost:6379
JWT_SECRET=generate-with-node-crypto-randomBytes-64-hex
CORS_ORIGIN=http://localhost:3000

frontend/.env.local

NEXT_PUBLIC_API_BASE=http://localhost:4000

Production Deployment

ServicePlatformNotes
BackendRender Web ServiceAuto-deploys from main
PostgreSQLRender PostgresFree tier — export before 90-day limit
RedisRender RedisFree tier — no persistence
FrontendVercelAuto-deploys from main, always-on

Deploy your own

  1. Fork this repo
  2. Create Render Postgres + Redis in Singapore region
  3. Create Render Web Service → root directory backend → runtime Node → add all env vars
  4. Run node src/db/init.js from Render Shell to apply schema
  5. Import repo to Vercel → root directory frontend → add NEXT_PUBLIC_API_BASE
  6. Update CORS_ORIGIN in Render env to your Vercel URL

Current Status

FeatureStatus
Docker Engine API metric collection✅
HTTP snapshot ingestion (no MQTT for users)✅
JWT auth + API key auth✅
Workspace isolation (multi-user)✅
Twin engine with health score✅
Rules-based alert engine✅
Decision engine with cooldown✅
Redis command queue (workspace-scoped)✅
Agent-pull self-healing execution✅
Full action lifecycle audit trail✅
What-if load simulation✅
Input validation + rate limiting✅
Jest test suite (9 tests)✅
GitHub Actions CI✅
Render + Vercel deployment✅
WebSocket live push❌ Polling every 5s
Kubernetes agent❌ Planned
ML anomaly detection❌ Planned

Planned Extensions

  • Kubernetes agent — watch Pod resources via K8s API, kubectl rollout restart as the action type
  • ML anomaly detection — Isolation Forest per twin, replaces fixed thresholds with learned baselines
  • WebSocket live push — Redis pub/sub → broadcaster, eliminates 5s polling lag
  • Scale-replica action — docker compose up --scale alongside restart

License

MIT — see LICENSE.


Built by Aaryan Bairagi