Phase 0: architecture docs, ADRs, and repo scaffold
Establish the design foundation for the heleosv2 multi-tenant hosting platform before any implementation code: - Monorepo skeleton: docs/, platform-infra/, site-templates/, deployments/, control-panel/ with orientation READMEs. - docs/: roadmap index, architecture + threat model, naming conventions, site profiles, provisioning workflow, backup & DR runbook, repo/GitOps layout, and the approved architecture plan. - docs/adr/: 9 ADRs recording the rationale for single-host Compose, Traefik edge, nginx+fpm split, shared MariaDB, ZFS-per-customer, decoupled backup streams, Forgejo, CLI-first, and SFTP-only. - Secrets hygiene: .gitignore (only *.enc.* committed) and .gitattributes (LF for scripts/Dockerfiles/YAML run on the Linux host). Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
This commit is contained in:
commit
9f819df4d9
25 changed files with 1223 additions and 0 deletions
23
.gitattributes
vendored
Normal file
23
.gitattributes
vendored
Normal file
|
|
@ -0,0 +1,23 @@
|
||||||
|
# Normalize line endings. This repo is authored on Windows but runs on a Linux
|
||||||
|
# host — CRLF in shell scripts / Dockerfiles / YAML breaks execution there.
|
||||||
|
|
||||||
|
# Default: let Git normalize text files to LF in the repo.
|
||||||
|
* text=auto eol=lf
|
||||||
|
|
||||||
|
# Files that MUST be LF (executed or parsed on Linux):
|
||||||
|
*.sh text eol=lf
|
||||||
|
*.bash text eol=lf
|
||||||
|
Dockerfile text eol=lf
|
||||||
|
*.dockerfile text eol=lf
|
||||||
|
*.yml text eol=lf
|
||||||
|
*.yaml text eol=lf
|
||||||
|
*.conf text eol=lf
|
||||||
|
*.env text eol=lf
|
||||||
|
*.py text eol=lf
|
||||||
|
*.j2 text eol=lf
|
||||||
|
|
||||||
|
# Binary files Git should never touch:
|
||||||
|
*.png binary
|
||||||
|
*.jpg binary
|
||||||
|
*.gz binary
|
||||||
|
*.age binary
|
||||||
33
.gitignore
vendored
Normal file
33
.gitignore
vendored
Normal file
|
|
@ -0,0 +1,33 @@
|
||||||
|
# Secrets — never commit plaintext. Use SOPS/age (*.enc.yaml is allowed).
|
||||||
|
*.env
|
||||||
|
.env
|
||||||
|
.env.*
|
||||||
|
!.env.example
|
||||||
|
secrets/
|
||||||
|
*.key
|
||||||
|
*.pem
|
||||||
|
!*.pub
|
||||||
|
|
||||||
|
# SOPS-encrypted files ARE allowed to be committed:
|
||||||
|
!*.enc.*
|
||||||
|
!*.sops.*
|
||||||
|
|
||||||
|
# Rendered runtime state that should not be tracked
|
||||||
|
deployments/**/.state/
|
||||||
|
deployments/**/data/
|
||||||
|
|
||||||
|
# OS / editor cruft
|
||||||
|
.DS_Store
|
||||||
|
Thumbs.db
|
||||||
|
*.swp
|
||||||
|
.idea/
|
||||||
|
.vscode/
|
||||||
|
|
||||||
|
# Ansible
|
||||||
|
*.retry
|
||||||
|
|
||||||
|
# Python (control-panel CLI)
|
||||||
|
__pycache__/
|
||||||
|
*.pyc
|
||||||
|
.venv/
|
||||||
|
venv/
|
||||||
39
README.md
Normal file
39
README.md
Normal file
|
|
@ -0,0 +1,39 @@
|
||||||
|
# heleosv2 — Multi-Tenant Web Hosting Platform
|
||||||
|
|
||||||
|
A rebuild of a small web-hosting business on modern, isolation-first
|
||||||
|
infrastructure: **container-per-customer + ZFS-dataset-per-customer**, with
|
||||||
|
repeatable CLI-driven provisioning and clean per-customer backup/restore.
|
||||||
|
|
||||||
|
> **Core philosophy — the customer is the boundary.** Isolation (containers),
|
||||||
|
> backup (ZFS), and restore all line up on the same boundary, so a compromised
|
||||||
|
> or broken site is contained and restorable without touching neighbours.
|
||||||
|
|
||||||
|
## Status
|
||||||
|
|
||||||
|
Phase 0 — writing the design documents before any platform code. See
|
||||||
|
[`docs/00-roadmap.md`](docs/00-roadmap.md).
|
||||||
|
|
||||||
|
## Repository layout
|
||||||
|
|
||||||
|
This monorepo groups four logical areas whose contents have different
|
||||||
|
lifecycles (see [`docs/07-repo-layout-gitops.md`](docs/07-repo-layout-gitops.md)):
|
||||||
|
|
||||||
|
| Path | Purpose |
|
||||||
|
|-------------------|---------------------------------------------------------------------|
|
||||||
|
| `docs/` | Architecture, ADRs, runbooks, conventions. |
|
||||||
|
| `platform-infra/` | Ansible + base compose for host, Traefik, MariaDB, Forgejo, monitoring. |
|
||||||
|
| `site-templates/` | Dockerfiles for standard images + compose templates per site profile. |
|
||||||
|
| `deployments/` | Rendered per-customer configs (GitOps state). **No plaintext secrets.** |
|
||||||
|
| `control-panel/` | Provisioning CLI now; customer-facing panel later. |
|
||||||
|
|
||||||
|
## Key decisions
|
||||||
|
|
||||||
|
- **Single bare-metal host**, Docker Compose per customer, no orchestrator.
|
||||||
|
- **Shared MariaDB** (per-site DB + least-privilege user).
|
||||||
|
- **Decoupled backups:** web files via ZFS snapshot/`send`; DB via
|
||||||
|
automysqlbackup + rsync; logs via Loki.
|
||||||
|
- **Lightweight DevOps:** Forgejo + built-in registry + Trivy.
|
||||||
|
- **CLI-first** provisioning; web panel deferred until the platform is proven.
|
||||||
|
|
||||||
|
The full approved architecture plan lives at
|
||||||
|
[`docs/architecture-plan.md`](docs/architecture-plan.md).
|
||||||
18
control-panel/README.md
Normal file
18
control-panel/README.md
Normal file
|
|
@ -0,0 +1,18 @@
|
||||||
|
# control-panel
|
||||||
|
|
||||||
|
The provisioning CLI (Phase 4), later the customer-facing web panel (Phase 8).
|
||||||
|
Empty until Phase 4.
|
||||||
|
|
||||||
|
**Planned CLI surface** (see
|
||||||
|
[../docs/05-provisioning-workflow.md](../docs/05-provisioning-workflow.md)):
|
||||||
|
|
||||||
|
| Command | Action |
|
||||||
|
|---------|--------|
|
||||||
|
| `provision` | Create a new site end-to-end (ZFS + DB + compose + route + SFTP). |
|
||||||
|
| `reconfigure` | Apply changes from an edited `site.yaml`. |
|
||||||
|
| `deprovision` | Tear down a site with gated, backed-up deletion. |
|
||||||
|
| `list` | Show all sites, profiles, status. |
|
||||||
|
| `backup` / `restore` | Ad-hoc backup; two-step (files + DB) restore. |
|
||||||
|
| `rotate-secret` | Regenerate DB/SFTP credentials. |
|
||||||
|
|
||||||
|
Idempotent, declarative (`site.yaml` is source of truth), logs to Loki.
|
||||||
17
deployments/README.md
Normal file
17
deployments/README.md
Normal file
|
|
@ -0,0 +1,17 @@
|
||||||
|
# deployments
|
||||||
|
|
||||||
|
GitOps state — one directory per site, rendered by the provisioning CLI
|
||||||
|
(Phase 4). Empty until the first site is provisioned.
|
||||||
|
|
||||||
|
Per-site layout (see
|
||||||
|
[../docs/03-naming-conventions.md](../docs/03-naming-conventions.md) §7):
|
||||||
|
|
||||||
|
```
|
||||||
|
<slug>/
|
||||||
|
├── site.yaml # declarative source of truth (slug, profile, domains)
|
||||||
|
├── docker-compose.yml # rendered from a site-templates profile
|
||||||
|
├── .env.example # non-secret references
|
||||||
|
└── secrets.enc.yaml # SOPS/age-encrypted secrets (committed encrypted only)
|
||||||
|
```
|
||||||
|
|
||||||
|
**Never commit plaintext secrets.** Only `*.enc.*` / `*.sops.*` are allowed.
|
||||||
46
docs/00-roadmap.md
Normal file
46
docs/00-roadmap.md
Normal file
|
|
@ -0,0 +1,46 @@
|
||||||
|
# heleosv2 — Roadmap & Document Index
|
||||||
|
|
||||||
|
This is the entry point for the platform design. Read the documents in order;
|
||||||
|
each builds on the previous.
|
||||||
|
|
||||||
|
## Document index
|
||||||
|
|
||||||
|
| # | Document | What it answers |
|
||||||
|
|---|----------|-----------------|
|
||||||
|
| 00 | **This file** | Phases, sequencing, where everything lives. |
|
||||||
|
| 01 | [Architecture & threat model](01-architecture-and-threat-model.md) | What the system is; which boundaries protect what. |
|
||||||
|
| 02 | [ADRs](adr/) | *Why* each major technology choice was made. |
|
||||||
|
| 03 | [Naming & conventions](03-naming-conventions.md) | How datasets, networks, containers, DBs, domains are named. |
|
||||||
|
| 04 | [Site profiles](04-site-profiles.md) | The four standard site stacks and their contents. |
|
||||||
|
| 05 | [Provisioning workflow](05-provisioning-workflow.md) | Signup → live, step by step. |
|
||||||
|
| 06 | [Backup & DR runbook](06-backup-and-dr.md) | How backups run and how to restore one customer. |
|
||||||
|
| 07 | [Repo layout & GitOps](07-repo-layout-gitops.md) | Repo boundaries, secrets, deployment flow. |
|
||||||
|
|
||||||
|
The full approved architecture plan is at
|
||||||
|
[`architecture-plan.md`](architecture-plan.md).
|
||||||
|
|
||||||
|
## Phased roadmap
|
||||||
|
|
||||||
|
| Phase | Goal | Output |
|
||||||
|
|-------|------|--------|
|
||||||
|
| **0** | Design docs & conventions | This `docs/` set (in progress). |
|
||||||
|
| **1** | Host baseline | Ansible: ZFS pool + datasets, Docker, nftables, SSH hardening, egress filtering. |
|
||||||
|
| **2** | Platform services | Traefik (TLS/ACME), shared MariaDB, Forgejo + registry. Smoke test: hello-world routed over HTTPS. |
|
||||||
|
| **3** | Site templates & images | Standard images (php-fpm non-root, nginx, static) + compose templates for 4 profiles; Trivy/gitleaks in CI. |
|
||||||
|
| **4** | Provisioning CLI | `provision` / `deprovision` / `list`; ZFS + DB + compose + route; per-customer chrooted SFTP. |
|
||||||
|
| **5** | Backup & DR | ZFS snapshot/`send` for web; automysqlbackup + rsync for DB; restore drill. |
|
||||||
|
| **6** | Observability | node_exporter + cAdvisor + Traefik metrics → Prometheus/Grafana; Loki/Promtail; Uptime-Kuma. |
|
||||||
|
| **7** | Migration | Move existing sites: static/redirect → custom PHP → WordPress; DNS cut per site. |
|
||||||
|
| **8** | Customer panel (later) | Web UI over the Phase 4 CLI; self-service backup/restore; Falco runtime detection. |
|
||||||
|
|
||||||
|
## Guiding principles
|
||||||
|
|
||||||
|
1. **The customer is the boundary** — isolation, backup, restore align on it.
|
||||||
|
2. **Containers isolate, they don't secure by themselves** — defense in depth
|
||||||
|
(non-root FPM, read-only rootfs, per-customer networks, egress filtering,
|
||||||
|
least-privilege DB users).
|
||||||
|
3. **Each backup stream matches its data's change pattern** — don't fold the DB
|
||||||
|
dump into the web dataset (it would bloat every incremental).
|
||||||
|
4. **CLI/templates before UI** — prove the platform, then wrap it.
|
||||||
|
5. **Phase the heavy stuff** — observability, panel, Harbor/Falco come after the
|
||||||
|
core works with real tenants.
|
||||||
124
docs/01-architecture-and-threat-model.md
Normal file
124
docs/01-architecture-and-threat-model.md
Normal file
|
|
@ -0,0 +1,124 @@
|
||||||
|
# 01 — Architecture & Threat Model
|
||||||
|
|
||||||
|
## 1. Purpose
|
||||||
|
|
||||||
|
Rebuild the hosting business as an **isolation-first, backup-clean,
|
||||||
|
repeatably-provisioned** platform on a single bare-metal host. Serves a mix of
|
||||||
|
WordPress, custom PHP, static HTML, and redirect-only sites.
|
||||||
|
|
||||||
|
## 2. High-level architecture
|
||||||
|
|
||||||
|
```
|
||||||
|
Internet
|
||||||
|
│
|
||||||
|
(80/443, SFTP, admin SSH)
|
||||||
|
│
|
||||||
|
┌────────▼────────┐
|
||||||
|
│ Traefik │ edge: TLS termination, ACME,
|
||||||
|
│ (edge router) │ dynamic label-based routing
|
||||||
|
└───┬─────────┬───┘
|
||||||
|
│ │ shared "proxy" network
|
||||||
|
┌───────────────┘ └───────────────┐
|
||||||
|
│ per-customer network A │ per-customer network B
|
||||||
|
┌────▼─────┐ ┌──────────┐ ┌─────▼────┐ ┌──────────┐
|
||||||
|
│ nginx │──▶│ php-fpm │ │ nginx │──▶│ php-fpm │
|
||||||
|
│ (site A) │ │ (site A) │ │ (site B) │ │ (site B) │
|
||||||
|
└────┬─────┘ └────┬─────┘ └────┬─────┘ └────┬─────┘
|
||||||
|
│ web root vol │ │ web root vol │
|
||||||
|
│ └──────────┐ ┌──────────┘ │
|
||||||
|
│ ▼ ▼ │
|
||||||
|
│ ┌─────────────────┐ │
|
||||||
|
│ │ shared MariaDB │ per-site DB + │
|
||||||
|
│ │ (platform net) │ least-priv user │
|
||||||
|
│ └─────────────────┘ │
|
||||||
|
▼ ▼
|
||||||
|
ZFS tank/customers/A/web ZFS tank/customers/B/web
|
||||||
|
```
|
||||||
|
|
||||||
|
Platform services (not shown per-customer): Forgejo + registry, Prometheus /
|
||||||
|
Grafana / Loki, Uptime-Kuma. All run as their own compose projects on the host.
|
||||||
|
|
||||||
|
### Request path
|
||||||
|
1. DNS points the customer domain at the host.
|
||||||
|
2. Traefik terminates TLS (Let's Encrypt) and routes by Host rule (from compose
|
||||||
|
labels) onto the customer's network.
|
||||||
|
3. For PHP profiles: nginx serves static assets and proxies `.php` over FastCGI
|
||||||
|
to that site's php-fpm. For static/redirect: Traefik or a tiny nginx answers
|
||||||
|
directly.
|
||||||
|
4. php-fpm talks to shared MariaDB over the platform network using the site's
|
||||||
|
own database + least-privilege credentials.
|
||||||
|
|
||||||
|
### Why Traefik *and* nginx (not either/or)
|
||||||
|
Traefik does not speak FastCGI, so it cannot talk to php-fpm directly. Traefik
|
||||||
|
owns the edge (TLS, routing, ACME); a small per-site nginx bridges HTTP→FastCGI.
|
||||||
|
This is deliberate, not redundancy.
|
||||||
|
|
||||||
|
## 3. Isolation boundaries
|
||||||
|
|
||||||
|
| Boundary | Mechanism | Protects against |
|
||||||
|
|----------|-----------|------------------|
|
||||||
|
| Process/filesystem | Separate containers per site | One site reading another's files/processes |
|
||||||
|
| Network (east-west) | Per-customer Docker bridge; only Traefik bridges to `proxy` | Site A reaching Site B's containers |
|
||||||
|
| Data at rest | ZFS dataset per customer, bind-mounted web root | Cross-customer data access; enables clean restore |
|
||||||
|
| Database | Per-site DB + least-privilege user on shared instance | Site A reading Site B's tables |
|
||||||
|
| Privilege | php-fpm non-root; read-only rootfs where possible; writable web root only | Privilege escalation within a container |
|
||||||
|
| Egress | Firewall egress filtering (esp. SMTP) | Hacked site sending spam / exfiltration |
|
||||||
|
|
||||||
|
## 4. Threat model
|
||||||
|
|
||||||
|
### 4.1 Assets
|
||||||
|
- Customer website files and databases.
|
||||||
|
- The host itself (kernel, Docker daemon, ZFS pool).
|
||||||
|
- Platform credentials (DB root, Traefik/ACME, Forgejo, SSH).
|
||||||
|
- Backups (onsite snapshots + offsite copies).
|
||||||
|
|
||||||
|
### 4.2 Primary threat actors & scenarios
|
||||||
|
1. **Compromised WordPress/PHP app** (most likely). Attacker gets code execution
|
||||||
|
inside one site's php-fpm container.
|
||||||
|
2. **Malicious/abusive tenant.**
|
||||||
|
3. **Credential theft** (leaked DB or SFTP creds).
|
||||||
|
4. **External network attacker** probing exposed ports.
|
||||||
|
|
||||||
|
### 4.3 What each scenario can and cannot do
|
||||||
|
|
||||||
|
**Compromised app container (S1):**
|
||||||
|
- ✅ Contained to: that site's files, that site's DB (its creds only), its own
|
||||||
|
network namespace.
|
||||||
|
- ❌ Blocked from: other customers' files (separate datasets/containers), other
|
||||||
|
DBs (least-privilege user), other customers' networks (no route), spamming
|
||||||
|
(egress SMTP filtered).
|
||||||
|
- ⚠️ **Residual risk:** containers share the host kernel — a kernel or Docker
|
||||||
|
escape breaks isolation. Mitigate with: patched host, non-root FPM, dropped
|
||||||
|
capabilities, no `--privileged`, read-only rootfs, seccomp defaults; later
|
||||||
|
gVisor and/or Falco runtime detection.
|
||||||
|
|
||||||
|
**Malicious tenant (S2):** same containment as S1, plus resource limits
|
||||||
|
(CPU/memory per compose project) to prevent noisy-neighbour DoS.
|
||||||
|
|
||||||
|
**Credential theft (S3):** blast radius limited to that one site because every
|
||||||
|
site has its own DB user and its own SFTP chroot. Rotate via CLI.
|
||||||
|
|
||||||
|
**Network attacker (S4):** only 80/443, SFTP, and admin SSH are exposed;
|
||||||
|
everything else denied by nftables. Admin SSH key-only + hardened.
|
||||||
|
|
||||||
|
### 4.4 Explicit non-goals / accepted risks
|
||||||
|
- **Containers are not a hardened sandbox.** We accept shared-kernel risk on a
|
||||||
|
single host and mitigate in depth rather than claiming VM-grade isolation.
|
||||||
|
- **Backups are crash-consistent, not transactionally atomic** across the
|
||||||
|
web-file and DB streams (see [06](06-backup-and-dr.md)). Acceptable for
|
||||||
|
PHP/WordPress workloads.
|
||||||
|
- No HA / multi-host failover in the initial design (single host by choice).
|
||||||
|
|
||||||
|
## 5. Security controls checklist (implemented across phases)
|
||||||
|
- [ ] Host: nftables default-deny inbound, egress SMTP filtered, SSH key-only.
|
||||||
|
- [ ] Docker: no `--privileged`, drop capabilities, read-only rootfs where
|
||||||
|
possible, per-project resource limits, userns considered.
|
||||||
|
- [ ] Per-customer network isolation; Traefik the only cross-network bridge.
|
||||||
|
- [ ] php-fpm runs as non-root; web root the only writable mount.
|
||||||
|
- [ ] Per-site DB users with least privilege; no shared DB accounts.
|
||||||
|
- [ ] SFTP chrooted per customer.
|
||||||
|
- [ ] Images scanned (Trivy) and secrets scanned (gitleaks) in CI.
|
||||||
|
- [ ] Automatic TLS via Traefik/ACME; HSTS.
|
||||||
|
- [ ] (Later) Falco runtime anomaly detection; gVisor for higher-risk tenants.
|
||||||
|
|
||||||
|
See [ADRs](adr/) for the rationale behind each major choice.
|
||||||
99
docs/03-naming-conventions.md
Normal file
99
docs/03-naming-conventions.md
Normal file
|
|
@ -0,0 +1,99 @@
|
||||||
|
# 03 — Naming & Conventions
|
||||||
|
|
||||||
|
Consistent, predictable names are what make CLI-driven provisioning and
|
||||||
|
scripted backup/restore reliable. Every resource for a site is derivable from a
|
||||||
|
single **slug**.
|
||||||
|
|
||||||
|
## 1. The slug
|
||||||
|
|
||||||
|
Each site has one canonical **slug**: lowercase, ASCII, `[a-z0-9-]`, 3–32 chars,
|
||||||
|
starting with a letter. Derived from the primary domain (dots → hyphens) or set
|
||||||
|
explicitly.
|
||||||
|
|
||||||
|
- `example.com` → `example-com`
|
||||||
|
- `blog.example.com` → `blog-example-com`
|
||||||
|
|
||||||
|
The slug is the join key across ZFS, Docker, database, and deployment config. It
|
||||||
|
never changes for the life of the site (renaming = new slug + migration).
|
||||||
|
|
||||||
|
> A customer may own several sites. Where a **customer** grouping is needed
|
||||||
|
> (billing, SFTP account), use a separate `customer` id with the same charset
|
||||||
|
> rules. The default is one slug per site.
|
||||||
|
|
||||||
|
## 2. ZFS datasets
|
||||||
|
|
||||||
|
```
|
||||||
|
tank/
|
||||||
|
├── customers/
|
||||||
|
│ └── <slug>/
|
||||||
|
│ └── web # web root ONLY (bind-mounted into the site)
|
||||||
|
└── platform/
|
||||||
|
├── docker # Docker data-root (images/layers/volumes)
|
||||||
|
├── mariadb # shared MariaDB datadir
|
||||||
|
├── db-backups/<slug> # automysqlbackup output per site
|
||||||
|
├── traefik # ACME store + dynamic config
|
||||||
|
├── forgejo # Git + registry data
|
||||||
|
└── monitoring # Prometheus/Loki data
|
||||||
|
```
|
||||||
|
|
||||||
|
Snapshots: `tank/customers/<slug>/web@auto-YYYYMMDD-HHMM`.
|
||||||
|
|
||||||
|
## 3. Docker resources
|
||||||
|
|
||||||
|
| Resource | Pattern | Example |
|
||||||
|
|----------|---------|---------|
|
||||||
|
| Compose project | `<slug>` | `example-com` |
|
||||||
|
| Container | `<slug>_<role>` | `example-com_nginx`, `example-com_fpm` |
|
||||||
|
| Per-customer network | `<slug>_net` | `example-com_net` |
|
||||||
|
| Shared edge network | `proxy` (fixed) | `proxy` |
|
||||||
|
| Platform network | `platform` (fixed) | `platform` (MariaDB, monitoring) |
|
||||||
|
| Named volume (if used) | `<slug>_<purpose>` | `example-com_fpmtmp` |
|
||||||
|
|
||||||
|
Traefik router/service labels also key off the slug:
|
||||||
|
`traefik.http.routers.<slug>.rule=Host(...)`.
|
||||||
|
|
||||||
|
## 4. Database (shared MariaDB)
|
||||||
|
|
||||||
|
| Resource | Pattern | Example |
|
||||||
|
|----------|---------|---------|
|
||||||
|
| Database | `db_<slug_underscored>` | `db_example_com` |
|
||||||
|
| DB user | `u_<slug_underscored>` | `u_example_com` |
|
||||||
|
| Grants | `ALL PRIVILEGES ON db_<...>.* ` to that user only | — |
|
||||||
|
|
||||||
|
Hyphens in the slug become underscores for MySQL identifiers
|
||||||
|
(`example-com` → `example_com`). Passwords are generated, stored encrypted
|
||||||
|
(SOPS/age), never reused across sites.
|
||||||
|
|
||||||
|
## 5. Domains & TLS
|
||||||
|
|
||||||
|
- Primary domain drives the slug; additional aliases are listed in the site's
|
||||||
|
deployment config and added to the Traefik Host rule.
|
||||||
|
- Certificates are issued per domain automatically by Traefik/ACME; no manual
|
||||||
|
naming.
|
||||||
|
|
||||||
|
## 6. SFTP accounts
|
||||||
|
|
||||||
|
| Resource | Pattern | Example |
|
||||||
|
|----------|---------|---------|
|
||||||
|
| SFTP user | `sftp_<slug_underscored>` | `sftp_example_com` |
|
||||||
|
| Chroot path | `tank/customers/<slug>/web` | — |
|
||||||
|
|
||||||
|
## 7. Deployment config (in `deployments/`)
|
||||||
|
|
||||||
|
```
|
||||||
|
deployments/
|
||||||
|
└── <slug>/
|
||||||
|
├── docker-compose.yml # rendered from a site-templates profile
|
||||||
|
├── .env.example # non-secret defaults / references
|
||||||
|
├── secrets.enc.yaml # SOPS/age-encrypted secrets (committed)
|
||||||
|
└── site.yaml # slug, profile, domains, options (source of truth)
|
||||||
|
```
|
||||||
|
|
||||||
|
`site.yaml` is the declarative source of truth the CLI reads/writes; everything
|
||||||
|
else is rendered from it.
|
||||||
|
|
||||||
|
## 8. Reserved words
|
||||||
|
|
||||||
|
Slugs may not be: `proxy`, `platform`, `traefik`, `mariadb`, `forgejo`,
|
||||||
|
`monitoring`, `docker`, `customers`, `platform-infra`, `site-templates`,
|
||||||
|
`deployments`, `control-panel` (avoids collisions with platform names).
|
||||||
91
docs/04-site-profiles.md
Normal file
91
docs/04-site-profiles.md
Normal file
|
|
@ -0,0 +1,91 @@
|
||||||
|
# 04 — Site Profiles
|
||||||
|
|
||||||
|
A **profile** is a standard site stack with a compose template in
|
||||||
|
`site-templates/`. The provisioning CLI renders a profile with a site's slug,
|
||||||
|
domains, and options. Four profiles cover all current workloads.
|
||||||
|
|
||||||
|
Common to every profile:
|
||||||
|
- Traefik labels for routing + TLS (Host rule from domains, ACME resolver).
|
||||||
|
- Joined to the site's own `<slug>_net` network; only Traefik bridges to `proxy`.
|
||||||
|
- Web root bind-mounted from `tank/customers/<slug>/web`.
|
||||||
|
- Per-project CPU/memory limits.
|
||||||
|
- Containers run non-root; rootfs read-only where the profile allows.
|
||||||
|
|
||||||
|
| Profile | Web server | PHP | Database | Typical use |
|
||||||
|
|---------|-----------|-----|----------|-------------|
|
||||||
|
| `static` | nginx (or Traefik direct) | — | — | HTML/JS sites, landing pages |
|
||||||
|
| `redirect` | Traefik router / tiny nginx | — | — | Domain redirects |
|
||||||
|
| `custom-php` | nginx | php-fpm | optional | Bespoke PHP apps |
|
||||||
|
| `wordpress` | nginx | php-fpm | required | WordPress sites |
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## `static`
|
||||||
|
|
||||||
|
- **Containers:** one nginx serving the web root read-only. For very simple
|
||||||
|
cases, Traefik can serve files directly with no container.
|
||||||
|
- **Volumes:** `tank/customers/<slug>/web` → `/usr/share/nginx/html` (read-only).
|
||||||
|
- **DB:** none.
|
||||||
|
- **Notes:** cheapest profile; near-zero backup incrementals when unchanged.
|
||||||
|
|
||||||
|
## `redirect`
|
||||||
|
|
||||||
|
- **Containers:** none preferred — implemented as a Traefik router rule with a
|
||||||
|
redirect middleware (`RedirectRegex`/`RedirectScheme`). A tiny nginx is the
|
||||||
|
fallback if complex rewrite logic is needed.
|
||||||
|
- **Volumes:** none.
|
||||||
|
- **DB:** none.
|
||||||
|
- **Config:** target URL(s) and redirect type (301/302), preserve-path flag —
|
||||||
|
all in `site.yaml`.
|
||||||
|
|
||||||
|
## `custom-php`
|
||||||
|
|
||||||
|
- **Containers:**
|
||||||
|
- `nginx` — serves static assets, proxies `.php` to fpm over FastCGI.
|
||||||
|
- `php-fpm` — non-root; only the web root (and a small tmp) writable.
|
||||||
|
- **Volumes:** `tank/customers/<slug>/web` shared by both (web root).
|
||||||
|
- **DB:** optional — if requested, a `db_<slug>` + `u_<slug>` on shared MariaDB;
|
||||||
|
credentials injected via env from `secrets.enc.yaml`.
|
||||||
|
- **Options:** PHP version (pinned image tag), extensions, `php.ini` overrides,
|
||||||
|
cron (via a scheduled fpm exec) if needed.
|
||||||
|
|
||||||
|
## `wordpress`
|
||||||
|
|
||||||
|
- **Containers:** same nginx + php-fpm pair as `custom-php`, using a WordPress
|
||||||
|
base image (or WP installed into the web root on first provision).
|
||||||
|
- **Volumes:** `tank/customers/<slug>/web` (WordPress core, themes, plugins,
|
||||||
|
uploads).
|
||||||
|
- **DB:** **required** — `db_<slug>` + `u_<slug>` on shared MariaDB.
|
||||||
|
- **Hardening baked in:**
|
||||||
|
- php-fpm non-root; `wp-content/uploads` writable, PHP execution denied there
|
||||||
|
(nginx rule) to blunt upload-based RCE.
|
||||||
|
- Read-only rootfs where WordPress tolerates it; XML-RPC restricted; sensible
|
||||||
|
security headers via Traefik middleware.
|
||||||
|
- WP-CLI available for provisioning/maintenance (install, update, search-
|
||||||
|
replace on domain migration).
|
||||||
|
- **Options:** PHP version, multisite flag, alias domains, initial admin user.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## Profile inputs (from `site.yaml`)
|
||||||
|
|
||||||
|
Every profile is rendered from the same declarative fields; unused fields are
|
||||||
|
ignored per profile:
|
||||||
|
|
||||||
|
```yaml
|
||||||
|
slug: example-com
|
||||||
|
profile: wordpress # static | redirect | custom-php | wordpress
|
||||||
|
domains: # first is primary → drives cert + slug
|
||||||
|
- example.com
|
||||||
|
- www.example.com
|
||||||
|
php_version: "8.3" # custom-php / wordpress
|
||||||
|
database: true # custom-php (wordpress forces true)
|
||||||
|
redirect_to: null # redirect profile only
|
||||||
|
resources:
|
||||||
|
cpu: "1.0"
|
||||||
|
memory: "512m"
|
||||||
|
```
|
||||||
|
|
||||||
|
Adding a profile = adding a template in `site-templates/` and a case in the CLI
|
||||||
|
renderer. Keep the set small; special cases are options on a profile, not new
|
||||||
|
profiles.
|
||||||
77
docs/05-provisioning-workflow.md
Normal file
77
docs/05-provisioning-workflow.md
Normal file
|
|
@ -0,0 +1,77 @@
|
||||||
|
# 05 — Provisioning Workflow
|
||||||
|
|
||||||
|
How a site goes from request to live, and back off again. All steps are
|
||||||
|
operator-driven via the `control-panel` CLI in Phase 0–7; the customer panel
|
||||||
|
(Phase 8) calls the same operations.
|
||||||
|
|
||||||
|
## Design principle
|
||||||
|
|
||||||
|
The CLI is **declarative and idempotent**: `site.yaml` is the source of truth,
|
||||||
|
everything else is rendered from it, and re-running a command converges to the
|
||||||
|
desired state rather than erroring. Each step is individually re-runnable so a
|
||||||
|
failed provision can be resumed.
|
||||||
|
|
||||||
|
## New site — `provision`
|
||||||
|
|
||||||
|
Input: `slug`, `profile`, `domains`, profile options (see
|
||||||
|
[04](04-site-profiles.md) `site.yaml`).
|
||||||
|
|
||||||
|
1. **Validate** — slug charset/reserved-word check
|
||||||
|
([03](03-naming-conventions.md)), domains resolvable/owned, profile known.
|
||||||
|
2. **ZFS** — create `tank/customers/<slug>/web` (and `db-backups/<slug>` if the
|
||||||
|
profile uses a DB).
|
||||||
|
3. **Database** (DB profiles) — create `db_<slug>` + `u_<slug>` with least
|
||||||
|
privilege on shared MariaDB; generate password; write to `secrets.enc.yaml`
|
||||||
|
(SOPS/age).
|
||||||
|
4. **Render** — produce `deployments/<slug>/docker-compose.yml` + `.env` from the
|
||||||
|
profile template and `site.yaml`.
|
||||||
|
5. **SFTP** — create chrooted `sftp_<slug>` account bound to the web root.
|
||||||
|
6. **Commit** — commit `deployments/<slug>/` to the `deployments` repo (audit
|
||||||
|
trail; secrets committed only in encrypted form).
|
||||||
|
7. **Deploy** — `docker compose up -d` in `deployments/<slug>/`. Traefik
|
||||||
|
discovers the route from labels; ACME issues the certificate.
|
||||||
|
8. **Verify** — HTTPS reachability + valid cert; for WordPress, run WP-CLI
|
||||||
|
install. Print access details.
|
||||||
|
|
||||||
|
DNS: the operator ensures the domain points at the host (pre- or post-provision;
|
||||||
|
the cert completes once DNS resolves).
|
||||||
|
|
||||||
|
## Change a site — `reconfigure`
|
||||||
|
|
||||||
|
Edit `site.yaml` (e.g. add an alias domain, bump PHP version, adjust limits),
|
||||||
|
then re-run: the CLI re-renders, re-commits, and `compose up -d` applies the
|
||||||
|
delta. TLS for new domains is automatic.
|
||||||
|
|
||||||
|
## Remove a site — `deprovision`
|
||||||
|
|
||||||
|
1. `docker compose down` (optionally `--remove-orphans`).
|
||||||
|
2. **Final backup** — take a last ZFS snapshot + final DB dump, retained per the
|
||||||
|
deprovision retention policy before deletion.
|
||||||
|
3. Drop `db_<slug>` + `u_<slug>` (after the final dump).
|
||||||
|
4. Remove the SFTP account.
|
||||||
|
5. Destroy `tank/customers/<slug>/web` (and `db-backups/<slug>`) **after** the
|
||||||
|
retention window — never immediately.
|
||||||
|
6. Remove `deployments/<slug>/` and commit.
|
||||||
|
|
||||||
|
> Destroys are gated: the CLI refuses to delete data younger than the retention
|
||||||
|
> window without an explicit `--force`, and always snapshots before destroying.
|
||||||
|
|
||||||
|
## Backup / restore
|
||||||
|
|
||||||
|
Routine backups run on a schedule (not per-command); restore is on demand. Both
|
||||||
|
are specified in the [Backup & DR runbook](06-backup-and-dr.md). CLI surface:
|
||||||
|
`backup <slug>` (ad-hoc), `restore <slug> --snapshot <name> --db <dump>`.
|
||||||
|
|
||||||
|
## CLI command summary
|
||||||
|
|
||||||
|
| Command | Action |
|
||||||
|
|---------|--------|
|
||||||
|
| `provision` | Create a new site end-to-end (steps 1–8). |
|
||||||
|
| `reconfigure` | Apply changes from an edited `site.yaml`. |
|
||||||
|
| `deprovision` | Tear down a site with gated, backed-up deletion. |
|
||||||
|
| `list` | Show all sites, profiles, status. |
|
||||||
|
| `backup` | Ad-hoc snapshot + DB dump for a site. |
|
||||||
|
| `restore` | Two-step restore (files + DB) for a site. |
|
||||||
|
| `rotate-secret` | Regenerate DB/SFTP credentials for a site. |
|
||||||
|
|
||||||
|
Every command is idempotent and logs to the platform log stream (Loki).
|
||||||
80
docs/06-backup-and-dr.md
Normal file
80
docs/06-backup-and-dr.md
Normal file
|
|
@ -0,0 +1,80 @@
|
||||||
|
# 06 — Backup & Disaster Recovery Runbook
|
||||||
|
|
||||||
|
Backups are **three decoupled streams**, each matched to its data's change
|
||||||
|
pattern (see [ADR 0006](adr/0006-decoupled-backup-streams.md)). This keeps
|
||||||
|
offsite transfer cheap while preserving per-customer restore.
|
||||||
|
|
||||||
|
| Stream | Data | Mechanism | Onsite | Offsite |
|
||||||
|
|--------|------|-----------|--------|---------|
|
||||||
|
| A | Web files | ZFS snapshot + `zfs send` | snapshots on `tank` | incremental `send` to offsite pool |
|
||||||
|
| B | Databases | automysqlbackup (per-DB, rotated) | `tank/platform/db-backups/<slug>` | rsync to offsite |
|
||||||
|
| C | Logs | Promtail → Loki | Loki store | (per Loki retention; not customer-restore data) |
|
||||||
|
|
||||||
|
Platform state (Traefik ACME store, Forgejo data, `deployments` repo, Prometheus
|
||||||
|
config) is backed up with the same ZFS mechanism from `tank/platform/*`.
|
||||||
|
|
||||||
|
## 1. Schedule & retention
|
||||||
|
|
||||||
|
- **Stream A (files):** frequent snapshots (e.g. hourly + daily), `zfs send`
|
||||||
|
daily offsite. Retention via a snapshot-pruning policy (keep N hourly, N
|
||||||
|
daily, N weekly).
|
||||||
|
- **Stream B (DB):** automysqlbackup nightly with its native daily/weekly/monthly
|
||||||
|
rotation; rsync offsite right after.
|
||||||
|
- **Ordering (consistency):** run the nightly DB dump **immediately before** the
|
||||||
|
daily file snapshot so the two are as close in time as possible. The streams
|
||||||
|
are **crash-consistent, not atomic** — acceptable for PHP/WordPress.
|
||||||
|
|
||||||
|
## 2. Offsite targets
|
||||||
|
|
||||||
|
- **Files:** a second ZFS host/pool receiving incremental `zfs send` streams.
|
||||||
|
Incrementals are near-empty for unchanged sites — the reason logs/dumps are
|
||||||
|
kept out of the customer dataset.
|
||||||
|
- **DB dumps:** rsync to the same or another offsite location (delta transfer).
|
||||||
|
- Both offsite copies are the recovery source if the primary host is lost.
|
||||||
|
|
||||||
|
## 3. Restore — single customer (the critical drill)
|
||||||
|
|
||||||
|
Two coordinated steps:
|
||||||
|
|
||||||
|
**A. Web files (ZFS)**
|
||||||
|
1. Identify the target snapshot: `tank/customers/<slug>/web@auto-...`.
|
||||||
|
2. Restore by clone/rollback (or `zfs receive` from offsite if the host is
|
||||||
|
gone) into `tank/customers/<slug>/web`.
|
||||||
|
|
||||||
|
**B. Database (dump)**
|
||||||
|
3. Pick the matching dump from `db-backups/<slug>` (or offsite).
|
||||||
|
4. Recreate `db_<slug>` + `u_<slug>` if needed; import the dump into shared
|
||||||
|
MariaDB.
|
||||||
|
|
||||||
|
**C. Bring up**
|
||||||
|
5. `docker compose up -d` in `deployments/<slug>/`.
|
||||||
|
6. Verify site + data; for WordPress confirm site URL / run WP-CLI
|
||||||
|
`search-replace` if the domain changed.
|
||||||
|
|
||||||
|
CLI: `restore <slug> --snapshot <name> --db <dump-file>` wraps A–C.
|
||||||
|
|
||||||
|
## 4. Restore — full host (DR rebuild)
|
||||||
|
|
||||||
|
1. Provision a fresh host from `platform-infra` Ansible (Docker, ZFS, firewall).
|
||||||
|
2. `zfs receive` the platform datasets (Traefik/Forgejo/MariaDB datadir) and all
|
||||||
|
`tank/customers/*` from offsite.
|
||||||
|
3. Restore DB dumps as needed (or rely on the received MariaDB datadir, then
|
||||||
|
reconcile with latest dumps).
|
||||||
|
4. Bring platform services up (Traefik, MariaDB, Forgejo, monitoring).
|
||||||
|
5. `docker compose up -d` per site from the `deployments` repo.
|
||||||
|
6. Repoint DNS if the host address changed.
|
||||||
|
|
||||||
|
## 5. Verification drills (run regularly, not just once)
|
||||||
|
|
||||||
|
- [ ] **Single-customer restore** to a scratch location; site + data come back;
|
||||||
|
neighbours untouched.
|
||||||
|
- [ ] **Unchanged-site incremental is near-empty** — confirms the lean-dataset
|
||||||
|
goal ([ADR 0005](adr/0005-zfs-per-customer.md)) is actually holding.
|
||||||
|
- [ ] **DB dump imports cleanly** and row counts match expectations.
|
||||||
|
- [ ] **Offsite `zfs receive`** of a customer dataset succeeds on the DR target.
|
||||||
|
- [ ] **Full-host rebuild** rehearsed at least once end-to-end.
|
||||||
|
|
||||||
|
## 6. What is NOT backed up here
|
||||||
|
- Container images — rebuildable from `site-templates` + registry.
|
||||||
|
- Rendered runtime state (`.state/`, container-local `data/`) — regenerated.
|
||||||
|
- Loki logs are operational telemetry, not customer-restore data.
|
||||||
83
docs/07-repo-layout-gitops.md
Normal file
83
docs/07-repo-layout-gitops.md
Normal file
|
|
@ -0,0 +1,83 @@
|
||||||
|
# 07 — Repository Layout & GitOps
|
||||||
|
|
||||||
|
## 1. Why a monorepo with clear boundaries
|
||||||
|
|
||||||
|
Platform code, image templates, per-customer deployment state, and secrets have
|
||||||
|
**different lifecycles and audiences**. We keep them in one repository (simple to
|
||||||
|
reason about for a small team) but as **strictly separated top-level areas**, so
|
||||||
|
they can be split into independent repos later without restructuring.
|
||||||
|
|
||||||
|
```
|
||||||
|
heleosv2/
|
||||||
|
├── docs/ # this design set (ADRs, runbooks, conventions)
|
||||||
|
├── platform-infra/ # Ansible + base compose for host & platform services
|
||||||
|
├── site-templates/ # Dockerfiles + compose templates per site profile
|
||||||
|
├── deployments/ # rendered per-customer configs (GitOps state)
|
||||||
|
└── control-panel/ # provisioning CLI now; customer panel later
|
||||||
|
```
|
||||||
|
|
||||||
|
### `platform-infra/`
|
||||||
|
Host baseline and platform services as code:
|
||||||
|
- Ansible roles: ZFS pool/datasets, Docker, nftables (Docker-aware), SSH
|
||||||
|
hardening, egress filtering, automysqlbackup, ZFS snapshot/`send` jobs.
|
||||||
|
- Base compose projects: Traefik, shared MariaDB, Forgejo + registry,
|
||||||
|
Prometheus/Grafana/Loki, Uptime-Kuma.
|
||||||
|
|
||||||
|
### `site-templates/`
|
||||||
|
The building blocks the CLI renders from:
|
||||||
|
- Dockerfiles for standard images (php-fpm non-root, nginx, static base).
|
||||||
|
- One compose **template** per profile (`static`, `redirect`, `custom-php`,
|
||||||
|
`wordpress`) with placeholders filled from `site.yaml`.
|
||||||
|
- Image builds run through CI with Trivy + gitleaks; images pushed to the
|
||||||
|
Forgejo registry.
|
||||||
|
|
||||||
|
### `deployments/`
|
||||||
|
The **GitOps state** — one directory per site (see
|
||||||
|
[03](03-naming-conventions.md) §7). Rendered `docker-compose.yml`, `.env`,
|
||||||
|
encrypted `secrets.enc.yaml`, and the declarative `site.yaml`. Committing here is
|
||||||
|
the audit trail of what is deployed. **Never commit plaintext secrets.**
|
||||||
|
|
||||||
|
### `control-panel/`
|
||||||
|
The provisioning CLI (`provision`/`reconfigure`/`deprovision`/`backup`/`restore`/
|
||||||
|
…). Reads/writes `site.yaml`, renders from `site-templates/`, writes to
|
||||||
|
`deployments/`, and drives ZFS/DB/Docker. The future web panel lives here too,
|
||||||
|
calling the same operations.
|
||||||
|
|
||||||
|
## 2. Secrets
|
||||||
|
|
||||||
|
- **SOPS + age** encrypt secrets at rest; only `*.enc.*` / `*.sops.*` files are
|
||||||
|
committed (enforced by [`.gitignore`](../.gitignore)).
|
||||||
|
- Plaintext `.env` files are git-ignored; a decrypt step materializes runtime
|
||||||
|
env just before `compose up` (and it stays out of the customer ZFS dataset).
|
||||||
|
- Per-site DB/SFTP credentials are unique and rotatable (`rotate-secret`).
|
||||||
|
- The age private key is an operator secret, stored outside the repo and part of
|
||||||
|
DR (without it, encrypted secrets are unrecoverable — back it up offline).
|
||||||
|
|
||||||
|
## 3. GitOps flow
|
||||||
|
|
||||||
|
```
|
||||||
|
edit site.yaml ─▶ CLI renders ─▶ commit deployments/<slug> ─▶ compose up -d
|
||||||
|
(intent) (from templates) (audit trail) (converge)
|
||||||
|
```
|
||||||
|
|
||||||
|
- **Source of truth:** `site.yaml` per site + the base compose in
|
||||||
|
`platform-infra`.
|
||||||
|
- **Change =** a commit in `deployments/`. History shows who deployed what, when.
|
||||||
|
- **CI** (Forgejo Actions): lint/scan templates and images; optionally validate
|
||||||
|
that `deployments/` renders cleanly from `site.yaml`. Deployment stays
|
||||||
|
operator-triggered on the single host initially (no auto-apply agent yet).
|
||||||
|
|
||||||
|
## 4. Branching
|
||||||
|
|
||||||
|
- `main` is deployable. Platform/template changes go via short-lived branches +
|
||||||
|
PR + CI (scans must pass).
|
||||||
|
- `deployments/` commits may be direct on `main` (operational changes) but still
|
||||||
|
run the render/scan checks.
|
||||||
|
|
||||||
|
## 5. When to split into multiple repos
|
||||||
|
|
||||||
|
Split when any becomes true: multiple operators needing different access to
|
||||||
|
`deployments/` vs platform code; `deployments/` history dominating the repo; or
|
||||||
|
open-sourcing `site-templates`/`control-panel` while keeping `deployments`
|
||||||
|
private. The top-level separation above makes that a clean `git filter-repo`
|
||||||
|
extraction rather than a rewrite.
|
||||||
22
docs/adr/0001-single-host-docker-compose.md
Normal file
22
docs/adr/0001-single-host-docker-compose.md
Normal file
|
|
@ -0,0 +1,22 @@
|
||||||
|
# ADR 0001 — Single bare-metal host, Docker Compose per customer
|
||||||
|
|
||||||
|
**Status:** Accepted
|
||||||
|
|
||||||
|
## Context
|
||||||
|
Small hosting business, modest site count, ZFS-based backup strategy. Options
|
||||||
|
ranged from a single host to a multi-host orchestrated fleet (Swarm/Nomad/k8s)
|
||||||
|
or cloud VMs.
|
||||||
|
|
||||||
|
## Decision
|
||||||
|
Run everything on **one bare-metal host**, using **one Docker Compose project
|
||||||
|
per customer/site**. No orchestrator. Host configured with Ansible.
|
||||||
|
|
||||||
|
## Consequences
|
||||||
|
- ✅ Simplest possible operational model; ZFS lives directly on local disks
|
||||||
|
(best snapshot/`send` story, no network-storage complications).
|
||||||
|
- ✅ Compose-per-customer maps 1:1 to the isolation and backup boundary.
|
||||||
|
- ✅ No control-plane overhead (no etcd, no scheduler).
|
||||||
|
- ❌ No built-in HA/failover; the host is a single point of failure — mitigated
|
||||||
|
by disciplined offsite backups and a documented rebuild/restore drill.
|
||||||
|
- ❌ Vertical scaling only. If site count outgrows one box, revisit with a
|
||||||
|
superseding ADR (Swarm/Nomad, ZFS locality becomes the hard problem).
|
||||||
25
docs/adr/0002-traefik-as-edge.md
Normal file
25
docs/adr/0002-traefik-as-edge.md
Normal file
|
|
@ -0,0 +1,25 @@
|
||||||
|
# ADR 0002 — Traefik as the edge router
|
||||||
|
|
||||||
|
**Status:** Accepted
|
||||||
|
|
||||||
|
## Context
|
||||||
|
Multi-tenant host needs TLS termination, automatic certificates, and routing
|
||||||
|
that changes every time a site is added/removed. Alternatives: hand-managed
|
||||||
|
nginx vhosts, Caddy, HAProxy.
|
||||||
|
|
||||||
|
## Decision
|
||||||
|
Use **Traefik** as the single edge router: TLS termination, **Let's Encrypt
|
||||||
|
(ACME)** automation, and **dynamic label-based routing** driven by each
|
||||||
|
customer compose project's labels. One shared `proxy` Docker network connects
|
||||||
|
Traefik to each customer network.
|
||||||
|
|
||||||
|
## Consequences
|
||||||
|
- ✅ Adding a site needs no central config edit — Traefik discovers routes from
|
||||||
|
the new project's labels.
|
||||||
|
- ✅ Certificates are automatic and auto-renewed.
|
||||||
|
- ✅ Customer containers stay unreachable except through Traefik (it is the only
|
||||||
|
service bridging `proxy` and a customer network).
|
||||||
|
- ❌ Traefik **cannot speak FastCGI**, so PHP sites still need a per-site web
|
||||||
|
server — see [ADR 0003](0003-nginx-fpm-per-site.md).
|
||||||
|
- ⚠️ Traefik is now a critical single component; its config, ACME store, and
|
||||||
|
dashboard must be secured and backed up.
|
||||||
23
docs/adr/0003-nginx-fpm-per-site.md
Normal file
23
docs/adr/0003-nginx-fpm-per-site.md
Normal file
|
|
@ -0,0 +1,23 @@
|
||||||
|
# ADR 0003 — Per-site nginx + php-fpm behind Traefik
|
||||||
|
|
||||||
|
**Status:** Accepted
|
||||||
|
|
||||||
|
## Context
|
||||||
|
Traefik terminates TLS and routes but cannot talk FastCGI to php-fpm. PHP sites
|
||||||
|
need something that serves static assets and bridges HTTP→FastCGI. Options:
|
||||||
|
one combined nginx+fpm image, Caddy-with-`php_fastcgi` (single process), or the
|
||||||
|
classic split of nginx and php-fpm into two containers.
|
||||||
|
|
||||||
|
## Decision
|
||||||
|
Each PHP site runs **nginx + php-fpm as two containers** within its own compose
|
||||||
|
project, sharing the web-root volume. Traefik routes to the site's nginx; nginx
|
||||||
|
proxies `.php` to that site's php-fpm.
|
||||||
|
|
||||||
|
## Consequences
|
||||||
|
- ✅ Conventional, extremely well-documented pattern.
|
||||||
|
- ✅ PHP runtime can be patched/pinned per site independently of the web server.
|
||||||
|
- ✅ nginx cannot execute PHP outside the fpm boundary; php-fpm runs non-root.
|
||||||
|
- ❌ Two containers per PHP site increases container count and compose verbosity.
|
||||||
|
- ↔️ Caddy-single-container remains a valid future simplification for
|
||||||
|
low-traffic sites; revisit if container density becomes a concern. Static and
|
||||||
|
redirect profiles avoid php-fpm entirely.
|
||||||
27
docs/adr/0004-shared-mariadb.md
Normal file
27
docs/adr/0004-shared-mariadb.md
Normal file
|
|
@ -0,0 +1,27 @@
|
||||||
|
# ADR 0004 — Shared MariaDB instance, per-site DB + user
|
||||||
|
|
||||||
|
**Status:** Accepted
|
||||||
|
|
||||||
|
## Context
|
||||||
|
Two models: (a) a MariaDB container per customer — strong isolation, aligns with
|
||||||
|
per-customer ZFS snapshots, but ~100–300 MB idle RAM each; or (b) one shared
|
||||||
|
MariaDB instance with a separate database and least-privilege user per site —
|
||||||
|
RAM-efficient via a shared buffer pool, but the live datadir is shared.
|
||||||
|
|
||||||
|
## Decision
|
||||||
|
Run **one shared MariaDB** instance. Each site gets its **own database and its
|
||||||
|
own least-privilege user**. The instance's datadir lives on a dedicated platform
|
||||||
|
ZFS dataset, not inside any customer dataset.
|
||||||
|
|
||||||
|
## Consequences
|
||||||
|
- ✅ Efficient memory use (shared buffer pool) — favours density on one host.
|
||||||
|
- ✅ Per-site DB users mean a compromised site's credentials expose only that
|
||||||
|
site's database.
|
||||||
|
- ❌ The live datadir cannot be snapshotted per-customer atomically with their
|
||||||
|
files — addressed by decoupling DB backups (see
|
||||||
|
[ADR 0006](0006-decoupled-backup-streams.md)): per-database dumps via
|
||||||
|
automysqlbackup give per-customer restore granularity.
|
||||||
|
- ⚠️ The shared instance is a shared-fate component (a crash or bad query can
|
||||||
|
affect all sites) and a noisy-neighbour surface — mitigate with tuning and,
|
||||||
|
later, `mysqld_exporter` monitoring. Revisit per-customer DB if a tenant needs
|
||||||
|
strong isolation or a different engine/version.
|
||||||
25
docs/adr/0005-zfs-per-customer.md
Normal file
25
docs/adr/0005-zfs-per-customer.md
Normal file
|
|
@ -0,0 +1,25 @@
|
||||||
|
# ADR 0005 — ZFS dataset per customer
|
||||||
|
|
||||||
|
**Status:** Accepted
|
||||||
|
|
||||||
|
## Context
|
||||||
|
Backups must be cheap, integrity-checked, and restorable per customer without
|
||||||
|
touching neighbours. ZFS offers checksummed integrity, cheap copy-on-write
|
||||||
|
snapshots, and incremental `zfs send`.
|
||||||
|
|
||||||
|
## Decision
|
||||||
|
Create a **ZFS dataset per customer** (`tank/customers/<customer>/web`) holding
|
||||||
|
**only the web root**. Docker image/layer storage and the MariaDB datadir live
|
||||||
|
on separate platform datasets. Customer web data is bind-mounted into containers
|
||||||
|
from the customer dataset.
|
||||||
|
|
||||||
|
## Consequences
|
||||||
|
- ✅ Snapshot and `zfs send` operate at the customer granularity — restore one
|
||||||
|
customer independently.
|
||||||
|
- ✅ Data integrity via checksums; cheap frequent snapshots.
|
||||||
|
- ✅ Keeping the dataset to web files **only** (no logs, no DB dumps) means an
|
||||||
|
unchanged site produces a near-empty incremental, keeping offsite `send` cheap
|
||||||
|
— the explicit reason logs and DB dumps are stored elsewhere
|
||||||
|
([ADR 0006](0006-decoupled-backup-streams.md)).
|
||||||
|
- ❌ Ties the platform to a ZFS-capable host (OpenZFS on Linux); not portable to
|
||||||
|
arbitrary cloud block storage without rework.
|
||||||
35
docs/adr/0006-decoupled-backup-streams.md
Normal file
35
docs/adr/0006-decoupled-backup-streams.md
Normal file
|
|
@ -0,0 +1,35 @@
|
||||||
|
# ADR 0006 — Decoupled backup streams (files vs DB vs logs)
|
||||||
|
|
||||||
|
**Status:** Accepted
|
||||||
|
|
||||||
|
## Context
|
||||||
|
An earlier proposal wrote per-database dumps *into* each customer's ZFS dataset
|
||||||
|
just before snapshotting, so one snapshot contained files + DB. Problem: the dump
|
||||||
|
changes every night even for a static site, so every incremental `zfs send`
|
||||||
|
would ship a fresh dump — defeating the cheap-incremental goal of
|
||||||
|
[ADR 0005](0005-zfs-per-customer.md). Logs in the dataset cause the same bloat.
|
||||||
|
|
||||||
|
## Decision
|
||||||
|
Back up each data type with the mechanism that fits its change pattern, in
|
||||||
|
**three decoupled streams**:
|
||||||
|
|
||||||
|
1. **Web files** → ZFS snapshot + `zfs send` offsite (truly incremental;
|
||||||
|
near-zero when unchanged).
|
||||||
|
2. **Databases** → **automysqlbackup** produces per-database dumps with
|
||||||
|
daily/weekly/monthly rotation into `tank/platform/db-backups/<customer>`, then
|
||||||
|
**rsync** offsite. Per-database dumps preserve per-customer restore
|
||||||
|
granularity.
|
||||||
|
3. **Logs** → shipped to **Loki** via Promtail; never stored in the customer
|
||||||
|
dataset.
|
||||||
|
|
||||||
|
## Consequences
|
||||||
|
- ✅ Unchanged sites cost almost nothing to back up offsite.
|
||||||
|
- ✅ Each stream is independently tunable (retention, cadence, target).
|
||||||
|
- ❌ Restoring a customer is a **two-step** operation (files from ZFS, DB from
|
||||||
|
the dump repo) rather than a single snapshot rollback — documented in the
|
||||||
|
[DR runbook](../06-backup-and-dr.md).
|
||||||
|
- ⚠️ The streams are **crash-consistent, not atomically consistent**. Schedule
|
||||||
|
the nightly DB dump close to the ZFS snapshot; the small window is harmless for
|
||||||
|
PHP/WordPress (files-on-disk + DB-rows).
|
||||||
|
- ↔️ Optional future: place `db-backups` on its own dataset and `zfs send` it for
|
||||||
|
checksummed/immutable offsite instead of rsync.
|
||||||
23
docs/adr/0007-forgejo-over-gitlab.md
Normal file
23
docs/adr/0007-forgejo-over-gitlab.md
Normal file
|
|
@ -0,0 +1,23 @@
|
||||||
|
# ADR 0007 — Forgejo over GitLab
|
||||||
|
|
||||||
|
**Status:** Accepted
|
||||||
|
|
||||||
|
## Context
|
||||||
|
The platform needs Git hosting, a container registry, and CI. GitLab CE bundles
|
||||||
|
all of this plus scanning but wants 4 GB+ RAM just for itself — heavy on a
|
||||||
|
single shared host. Forgejo (community fork of Gitea) is lightweight and ships a
|
||||||
|
built-in package/container registry and GitHub-Actions-compatible CI.
|
||||||
|
|
||||||
|
## Decision
|
||||||
|
Use **Forgejo** for Git + built-in **container registry**, with **Forgejo
|
||||||
|
Actions** (or Woodpecker CI) for pipelines. Security scanning via **Trivy** and
|
||||||
|
secret scanning via **gitleaks** run inside CI. No Harbor, no GitLab, no Jenkins.
|
||||||
|
|
||||||
|
## Consequences
|
||||||
|
- ✅ Low RAM footprint leaves resources for customer sites.
|
||||||
|
- ✅ One tool covers repos + registry; scanning is added as CI steps.
|
||||||
|
- ✅ Actions are GitHub-compatible, so pipelines are portable and familiar.
|
||||||
|
- ❌ Less batteries-included than GitLab (no built-in RBAC-heavy registry,
|
||||||
|
vulnerability dashboards, etc.) — acceptable at this scale.
|
||||||
|
- ↔️ If serious registry RBAC/signing/scanning is later needed, adopt **Harbor**
|
||||||
|
as a superseding ADR rather than migrating to GitLab.
|
||||||
25
docs/adr/0008-cli-first-panel-later.md
Normal file
25
docs/adr/0008-cli-first-panel-later.md
Normal file
|
|
@ -0,0 +1,25 @@
|
||||||
|
# ADR 0008 — CLI/templating first, customer panel later
|
||||||
|
|
||||||
|
**Status:** Accepted
|
||||||
|
|
||||||
|
## Context
|
||||||
|
A customer-facing control panel (self-service create/delete/backup/restore over
|
||||||
|
Docker + ZFS) is the single largest build — realistically months. Options: adopt
|
||||||
|
and extend an existing PaaS (Coolify/CapRover/Cloudron), build a bespoke panel
|
||||||
|
up front, or ship internal CLI tooling first and defer the UI.
|
||||||
|
|
||||||
|
## Decision
|
||||||
|
Build **provisioning as an internal CLI / templating tool first**
|
||||||
|
(`control-panel/`). The customer-facing web panel comes **later**, as a wrapper
|
||||||
|
over the same, proven CLI operations.
|
||||||
|
|
||||||
|
## Consequences
|
||||||
|
- ✅ Lowest initial risk; the provisioning logic (ZFS + DB + compose + routing)
|
||||||
|
is validated end-to-end before any UI investment.
|
||||||
|
- ✅ The CLI doubles as the automation surface the future panel and CI call into
|
||||||
|
— no throwaway work.
|
||||||
|
- ✅ Keeps early focus on correctness of isolation/backup rather than UX.
|
||||||
|
- ❌ No customer self-service until Phase 8; provisioning is operator-driven in
|
||||||
|
the meantime.
|
||||||
|
- ↔️ Adopting an existing PaaS remains an option for the panel layer if building
|
||||||
|
it proves too costly; the CLI would still own the ZFS/backup specifics.
|
||||||
21
docs/adr/0009-sftp-only.md
Normal file
21
docs/adr/0009-sftp-only.md
Normal file
|
|
@ -0,0 +1,21 @@
|
||||||
|
# ADR 0009 — SFTP only (no FTP)
|
||||||
|
|
||||||
|
**Status:** Accepted
|
||||||
|
|
||||||
|
## Context
|
||||||
|
Customers need file access to their web root. Legacy FTP is plaintext
|
||||||
|
(credentials and data in the clear). FTPS adds TLS but is firewall-hostile
|
||||||
|
(dynamic data ports). SFTP runs over SSH on a single port.
|
||||||
|
|
||||||
|
## Decision
|
||||||
|
Provide **SFTP only**, with each customer **chrooted to their own ZFS dataset**
|
||||||
|
(web root). No FTP, no FTPS. A web-based file manager (e.g. Filebrowser) may be
|
||||||
|
added later alongside the customer panel for non-technical users.
|
||||||
|
|
||||||
|
## Consequences
|
||||||
|
- ✅ Encrypted transport; single well-known port; firewall-friendly.
|
||||||
|
- ✅ Chroot enforces the same per-customer boundary as containers and datasets.
|
||||||
|
- ❌ Some legacy customer tooling only speaks FTP — those users migrate to SFTP
|
||||||
|
clients (documented in onboarding).
|
||||||
|
- ⚠️ SFTP access is a path into the web root; keys/passwords are per-customer and
|
||||||
|
rotatable, and the chroot prevents traversal to other customers or the host.
|
||||||
19
docs/adr/README.md
Normal file
19
docs/adr/README.md
Normal file
|
|
@ -0,0 +1,19 @@
|
||||||
|
# Architecture Decision Records
|
||||||
|
|
||||||
|
Short records capturing *why* each major choice was made. Format: Context →
|
||||||
|
Decision → Consequences. Status is one of Proposed / Accepted / Superseded.
|
||||||
|
|
||||||
|
| ADR | Decision | Status |
|
||||||
|
|-----|----------|--------|
|
||||||
|
| [0001](0001-single-host-docker-compose.md) | Single bare-metal host, Docker Compose per customer | Accepted |
|
||||||
|
| [0002](0002-traefik-as-edge.md) | Traefik as the edge router | Accepted |
|
||||||
|
| [0003](0003-nginx-fpm-per-site.md) | Per-site nginx + php-fpm behind Traefik | Accepted |
|
||||||
|
| [0004](0004-shared-mariadb.md) | Shared MariaDB instance, per-site DB + user | Accepted |
|
||||||
|
| [0005](0005-zfs-per-customer.md) | ZFS dataset per customer | Accepted |
|
||||||
|
| [0006](0006-decoupled-backup-streams.md) | Decoupled backup streams (files vs DB vs logs) | Accepted |
|
||||||
|
| [0007](0007-forgejo-over-gitlab.md) | Forgejo (lightweight) over GitLab | Accepted |
|
||||||
|
| [0008](0008-cli-first-panel-later.md) | CLI/templating first, customer panel later | Accepted |
|
||||||
|
| [0009](0009-sftp-only.md) | SFTP only (no FTP) | Accepted |
|
||||||
|
|
||||||
|
New ADRs are additive and numbered sequentially. To reverse a decision, add a
|
||||||
|
new ADR that supersedes the old one rather than editing history.
|
||||||
227
docs/architecture-plan.md
Normal file
227
docs/architecture-plan.md
Normal file
|
|
@ -0,0 +1,227 @@
|
||||||
|
# Rebuild: Multi-Tenant Web Hosting Platform (heleosv2)
|
||||||
|
|
||||||
|
## Context
|
||||||
|
|
||||||
|
An existing small web-hosting business runs nginx + PHP-FPM + MySQL on a shared
|
||||||
|
host, serving a mix of custom PHP, WordPress, static HTML, and redirect-only
|
||||||
|
sites. The goal is to rebuild on modern best practices with **isolation**,
|
||||||
|
**clean backups**, and **repeatable provisioning** as first principles.
|
||||||
|
|
||||||
|
Core philosophy: **the customer is the boundary.** Isolation (containers),
|
||||||
|
backup (ZFS dataset), and restore all line up on that same boundary, so a
|
||||||
|
compromised or broken site is contained and can be restored without touching
|
||||||
|
neighbours.
|
||||||
|
|
||||||
|
This document is an **architecture + phased task plan**, not a code change. No
|
||||||
|
application code exists yet (greenfield). Implementation happens step-by-step
|
||||||
|
after this plan is approved.
|
||||||
|
|
||||||
|
### Decisions locked with the user
|
||||||
|
- **Host topology:** single bare-metal host → Docker Compose per customer, no
|
||||||
|
orchestrator, Ansible for host config, Terraform not needed initially.
|
||||||
|
- **Control panel:** CLI/templating tooling first; customer-facing panel later
|
||||||
|
once the platform is proven.
|
||||||
|
- **Database:** one **shared MariaDB** instance (per-site DB + per-site user).
|
||||||
|
Density over per-instance isolation.
|
||||||
|
- **DevOps stack:** lightweight — **Forgejo** + built-in registry + Forgejo
|
||||||
|
Actions/Woodpecker + Trivy.
|
||||||
|
|
||||||
|
### Reframing carried into the design
|
||||||
|
Containers are an **isolation** boundary, not a hard security boundary (shared
|
||||||
|
kernel). The real protections come from: non-root PHP-FPM, read-only rootfs
|
||||||
|
where possible, per-customer Docker networks, egress filtering, least-privilege
|
||||||
|
DB users, and (later) runtime detection. Docker alone is not the security story.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## Target Architecture
|
||||||
|
|
||||||
|
### Edge & routing
|
||||||
|
- **Traefik** as the single edge: TLS termination, automatic Let's Encrypt
|
||||||
|
(ACME), dynamic label-based routing. One shared `proxy` Docker network.
|
||||||
|
- Traefik **cannot speak FastCGI**, so each PHP site runs its own small web
|
||||||
|
server behind Traefik:
|
||||||
|
|
||||||
|
`Traefik (TLS + routing) → per-site nginx → per-site php-fpm`
|
||||||
|
|
||||||
|
- Static/redirect sites are served directly by Traefik or a tiny nginx — no PHP.
|
||||||
|
|
||||||
|
### Per-customer stack (one Docker Compose project per customer/site)
|
||||||
|
Standard **site profiles**, each a compose template:
|
||||||
|
- **wordpress** — nginx + php-fpm (+ shared-DB credentials)
|
||||||
|
- **custom-php** — nginx + php-fpm (+ optional shared-DB credentials)
|
||||||
|
- **static** — nginx (or Traefik direct), no PHP, no DB
|
||||||
|
- **redirect** — Traefik router rule / tiny nginx, no PHP, no DB
|
||||||
|
|
||||||
|
Per-customer isolation:
|
||||||
|
- Own Docker bridge network; only Traefik joins both `proxy` and the customer
|
||||||
|
network, so customer containers are unreachable except through Traefik.
|
||||||
|
- PHP-FPM runs **non-root**; rootfs read-only where the profile allows; web
|
||||||
|
root is the only writable mount.
|
||||||
|
|
||||||
|
### Storage — ZFS
|
||||||
|
- Dataset layout, e.g.:
|
||||||
|
`tank/customers/<customer>/web` (web root **only** — kept lean) and
|
||||||
|
`tank/platform/{docker,mariadb,db-backups,traefik,forgejo,monitoring}`.
|
||||||
|
- **Docker images/layers** on their own platform dataset; **customer web data**
|
||||||
|
bind-mounted from per-customer datasets (clean per-customer snapshots).
|
||||||
|
- The customer dataset deliberately holds **only web files** — no logs, no DB
|
||||||
|
dumps — so an unchanged site produces a near-empty incremental and offsite
|
||||||
|
`zfs send` stays cheap.
|
||||||
|
- Snapshots per customer on a schedule; **`zfs send` offsite** for backup/DR.
|
||||||
|
|
||||||
|
### Database (shared MariaDB) — decoupled backup stream
|
||||||
|
- One MariaDB container; each site gets its own database + least-privilege user.
|
||||||
|
- Live datadir lives on a **platform** dataset (not inside any customer dataset).
|
||||||
|
- **DB backups are a separate stream from web files** (not folded into the
|
||||||
|
customer ZFS dataset — that would make every incremental `zfs send` ship a
|
||||||
|
fresh dump even for static sites):
|
||||||
|
- **automysqlbackup** produces per-database dumps with daily/weekly/monthly
|
||||||
|
rotation into `tank/platform/db-backups/<customer>` — per-customer restore
|
||||||
|
granularity for free.
|
||||||
|
- **rsync** those dumps offsite (lean; only changed dumps move). Optional
|
||||||
|
later: put db-backups on its own dataset and `zfs send` it if checksummed/
|
||||||
|
immutable offsite is wanted.
|
||||||
|
- **Consistency:** schedule the nightly DB dump close to the ZFS snapshot. The
|
||||||
|
two streams aren't atomically consistent, but the small window is harmless for
|
||||||
|
PHP/WordPress (files-on-disk + DB-rows).
|
||||||
|
|
||||||
|
### File access
|
||||||
|
- **SFTP only** (drop FTP — plaintext). Each customer chrooted to their dataset.
|
||||||
|
- Optional **web file manager** (e.g. Filebrowser) added with the panel later.
|
||||||
|
|
||||||
|
### Network security
|
||||||
|
- Host firewall via **nftables**. Note explicitly: Docker manipulates iptables
|
||||||
|
and bypasses ufw — firewall rules must account for Docker's chains.
|
||||||
|
- Only 80/443 (+ SFTP, + admin SSH) exposed. Inter-customer traffic denied.
|
||||||
|
- **Egress filtering:** block outbound SMTP except via an approved relay — stops
|
||||||
|
a hacked WordPress becoming a spam source (common real incident).
|
||||||
|
|
||||||
|
### Observability (phased)
|
||||||
|
- Phase in: node_exporter + cAdvisor + Traefik metrics → Prometheus + Grafana;
|
||||||
|
Loki + Promtail for logs. Add mysqld_exporter + Alertmanager later.
|
||||||
|
- Uptime-Kuma for cheap uptime/status pages.
|
||||||
|
|
||||||
|
### DevOps stack (lightweight)
|
||||||
|
- **Forgejo** (git) + built-in container **registry** + **Forgejo Actions** (or
|
||||||
|
Woodpecker) for CI.
|
||||||
|
- **Trivy** (images/fs/IaC/secrets) + **gitleaks** in CI. **Falco** later for
|
||||||
|
runtime detection of compromised containers. No Harbor, no GitLab, no Jenkins.
|
||||||
|
|
||||||
|
### Provisioning model (CLI-first, GitOps-lite)
|
||||||
|
A CLI/templating tool takes `(customer, site-profile, domain)` and:
|
||||||
|
1. Creates the ZFS dataset(s).
|
||||||
|
2. Renders `docker-compose.yml` + `.env` from the profile template.
|
||||||
|
3. Creates the shared-DB database + user (for profiles needing DB).
|
||||||
|
4. Commits rendered config to the `deployments` repo (audit trail).
|
||||||
|
5. Runs `docker compose up -d` in the customer directory.
|
||||||
|
Deprovision/backup/restore are additional CLI subcommands. The customer-facing
|
||||||
|
panel is a later wrapper over this same tooling.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## Repository Structure
|
||||||
|
|
||||||
|
Small number of repos (not a monorepo, not one-per-thing) because secrets and
|
||||||
|
per-customer state have a different lifecycle than platform code:
|
||||||
|
|
||||||
|
- **`platform-infra`** — Ansible (Docker, ZFS, nftables, host bootstrap), base
|
||||||
|
compose for Traefik / monitoring / Forgejo.
|
||||||
|
- **`site-templates`** — Dockerfiles for standard images (wordpress, php-fpm,
|
||||||
|
static) + compose templates per site profile.
|
||||||
|
- **`deployments`** — one directory per customer; rendered compose + env refs.
|
||||||
|
**Secrets via SOPS/age or a vault — never plaintext.**
|
||||||
|
- **`control-panel`** — the CLI tooling now; the customer panel later.
|
||||||
|
|
||||||
|
Solo/small-team may start as one monorepo with these as top-level dirs, but keep
|
||||||
|
the **deployments/secrets boundary clean from day one**.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## Documents to Create Before Code
|
||||||
|
|
||||||
|
1. **Architecture overview + threat model** — boundaries and what they do/don't
|
||||||
|
protect (this doc is the seed).
|
||||||
|
2. **ADRs** — one short record per major call: Traefik edge, shared MariaDB,
|
||||||
|
Forgejo-vs-GitLab, CLI-first-vs-panel. Captures the *why*.
|
||||||
|
3. **Naming & conventions** — dataset, network, container, database, domain
|
||||||
|
naming.
|
||||||
|
4. **Site profile spec** — exact contents of wordpress / custom-php / static /
|
||||||
|
redirect.
|
||||||
|
5. **Provisioning workflow** — signup → live, step by step.
|
||||||
|
6. **Backup & DR runbook** — two decoupled streams (web files via ZFS
|
||||||
|
snapshot/send; DB via automysqlbackup + rsync), cadence and near-simultaneous
|
||||||
|
scheduling, offsite targets, single-customer two-step restore drill.
|
||||||
|
7. **Repo layout / GitOps doc.**
|
||||||
|
8. **Roadmap / phased task plan** (section below).
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## Phased Task Roadmap
|
||||||
|
|
||||||
|
### Phase 0 — Documents & decisions
|
||||||
|
- Write the 8 documents above (start from this plan).
|
||||||
|
- Finalize naming conventions and dataset layout.
|
||||||
|
|
||||||
|
### Phase 1 — Host baseline (Ansible)
|
||||||
|
- Provision the bare-metal host: ZFS pool + datasets, Docker, nftables (with
|
||||||
|
Docker-aware rules), admin SSH hardening, egress filtering.
|
||||||
|
- `platform-infra` repo with idempotent Ansible.
|
||||||
|
|
||||||
|
### Phase 2 — Platform services
|
||||||
|
- Deploy Traefik (ACME/TLS, `proxy` network, dashboards secured).
|
||||||
|
- Deploy shared MariaDB (own datadir dataset, tuned buffer pool).
|
||||||
|
- Deploy Forgejo + registry.
|
||||||
|
- Smoke test: a hello-world container routed + TLS via Traefik.
|
||||||
|
|
||||||
|
### Phase 3 — Site templates & images
|
||||||
|
- Build standard images (php-fpm non-root, nginx, static) in `site-templates`.
|
||||||
|
- Trivy + gitleaks scanning in CI on those images.
|
||||||
|
- Author compose templates for the 4 site profiles.
|
||||||
|
|
||||||
|
### Phase 4 — Provisioning CLI
|
||||||
|
- `control-panel` CLI: provision / deprovision / list, rendering templates,
|
||||||
|
creating ZFS datasets, creating DB+user, committing to `deployments`,
|
||||||
|
`compose up`.
|
||||||
|
- SFTP access per customer (chrooted).
|
||||||
|
|
||||||
|
### Phase 5 — Backup & DR
|
||||||
|
- **Web files:** ZFS snapshot schedule + `zfs send` offsite (lean incrementals).
|
||||||
|
- **DB:** automysqlbackup (per-DB, rotated) → `db-backups` → rsync offsite,
|
||||||
|
scheduled close to the snapshot.
|
||||||
|
- **Restore drill:** restore one test customer end-to-end from offsite — files
|
||||||
|
from ZFS **and** DB from the dump repo (two coordinated steps).
|
||||||
|
|
||||||
|
### Phase 6 — Observability
|
||||||
|
- node_exporter + cAdvisor + Traefik metrics → Prometheus + Grafana; Loki +
|
||||||
|
Promtail. Uptime-Kuma. Basic alerts.
|
||||||
|
|
||||||
|
### Phase 7 — Migration
|
||||||
|
- Migrate existing sites profile-by-profile (static/redirect first, then custom
|
||||||
|
PHP, then WordPress). Cut DNS per site after validation.
|
||||||
|
|
||||||
|
### Phase 8 — Customer panel (later)
|
||||||
|
- Web panel wrapping the Phase 4 CLI: create/delete site, trigger backup,
|
||||||
|
self-service restore, web file manager. Optionally Falco for runtime detection.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## Verification / Validation
|
||||||
|
|
||||||
|
Because this is infrastructure, "done" is proven by drills, not unit tests:
|
||||||
|
|
||||||
|
1. **Routing/TLS:** provision a test `static` site → reachable over HTTPS via
|
||||||
|
Traefik with a valid Let's Encrypt cert.
|
||||||
|
2. **Isolation:** from inside customer A's container, confirm you **cannot**
|
||||||
|
reach customer B's containers/network or read B's files; confirm outbound
|
||||||
|
SMTP is blocked.
|
||||||
|
3. **DB least-privilege:** confirm customer A's DB user cannot see B's database.
|
||||||
|
4. **Provisioning round-trip:** CLI provisions a WordPress site from zero
|
||||||
|
(dataset + DB + compose + route) and it serves; then deprovision cleans up.
|
||||||
|
5. **Backup/restore drill (the critical one):** ZFS snapshot + `zfs send` of web
|
||||||
|
files, automysqlbackup + rsync of the DB, both offsite → restore that single
|
||||||
|
customer onto a clean target (files from ZFS, DB from the dump repo) and
|
||||||
|
verify site + data come back, neighbours untouched. Also verify an **unchanged
|
||||||
|
site yields a near-empty ZFS incremental** (confirms the lean-dataset goal).
|
||||||
|
6. **CI security gate:** Trivy/gitleaks fail a build on a known-vulnerable image
|
||||||
|
or planted secret.
|
||||||
11
platform-infra/README.md
Normal file
11
platform-infra/README.md
Normal file
|
|
@ -0,0 +1,11 @@
|
||||||
|
# platform-infra
|
||||||
|
|
||||||
|
Host baseline and platform services as code (Phase 1–2). Empty until then.
|
||||||
|
|
||||||
|
**Planned contents:**
|
||||||
|
- Ansible roles: ZFS pool/datasets, Docker, nftables (Docker-aware), SSH
|
||||||
|
hardening, egress filtering, automysqlbackup, ZFS snapshot/`send` jobs.
|
||||||
|
- Base compose projects: Traefik, shared MariaDB, Forgejo + registry,
|
||||||
|
Prometheus/Grafana/Loki, Uptime-Kuma.
|
||||||
|
|
||||||
|
See [../docs/07-repo-layout-gitops.md](../docs/07-repo-layout-gitops.md).
|
||||||
10
site-templates/README.md
Normal file
10
site-templates/README.md
Normal file
|
|
@ -0,0 +1,10 @@
|
||||||
|
# site-templates
|
||||||
|
|
||||||
|
Building blocks the provisioning CLI renders from (Phase 3). Empty until then.
|
||||||
|
|
||||||
|
**Planned contents:**
|
||||||
|
- Dockerfiles for standard images: php-fpm (non-root), nginx, static base.
|
||||||
|
- One compose template per profile: `static`, `redirect`, `custom-php`,
|
||||||
|
`wordpress` — placeholders filled from a site's `site.yaml`.
|
||||||
|
|
||||||
|
Profiles are specified in [../docs/04-site-profiles.md](../docs/04-site-profiles.md).
|
||||||
Loading…
Add table
Reference in a new issue