12 KiB
12 KiB
Backup-Automation Architektur
Überblick
Das System automatisiert Backups aller in Netbox gekennzeichneten Systeme mit:
- Borg Backup für deduplizierte, verschlüsselte Backups
- Kestra für Workflow-Orchestrierung
- Ansible für standardisierte Installation
- Rclone für S3-Replication
- Hetzner S3 für Offsite-Storage
Komponenten
1. Netbox Integration
- Quelle der Wahrheit: Netbox IPAM + Inventar
- Markierung: TAG
backup=trueauf Systemen - API: REST für dynamische Abfragen
- Update-Frequenz: Täglich
2. Kestra Orchestration
- Rolle: Zentrale Workflow-Engine
- Auslöser: Systemd Timer (02:00 Uhr täglich)
- Aufgaben:
- Netbox abfragen → Liste Backup-Systeme
- SSH-Keys prüfen/generieren
- Ansible Playbook starten
- Status erfassen
- Fehlerbehandlung: Retry-Logik, Notifications
3. Ansible Automation
-
Playbooks:
install-borg-client.yaml- Clients Setupconfigure-backup-server.yaml- Server Setupupdate-backup-config.yaml- Config Updates
-
Roles:
borg-client- Borg Installationbackup-server- Server + Rclone
-
Inventory: Dynamisch aus Netbox (via Kestra)
4. Borg Backup
- Format: Deduplicating, Encrypting, Compressing Backup
- Encryption: AES-256-CTR
- Compression: LZ4 (schnell) oder Zstd (besser)
- Retention: 7 tägl., 4 wöchentl., 12 monatl.
- Repository-Struktur:
/backup/repos/{hostname}/ ├── data/ ├── index.*/ ├── hints └── nonce
5. Backup-Server
- Rolle: Zentrale Sammelstelle aller Backups
- User:
backup:backupmit restricted SSH - Verzeichnis:
/backup/repos/{hostname} - SSH-Zugang: Nur für
borg serve
6. Docker-Spezifika
- Pre-Backup Hooks: MySQL/PostgreSQL Dumps
- Named Volumes: Automatisch exportiert
- Compose-Dateien: Backup als YAML
- Cleanup: Alte Dumps regelmäßig löschen
7. Rclone S3-Sync
- Quelle:
/backup/repos(Backup-Server) - Ziel:
hetzner:mbo-backups/production/repos/ - Strategie: Incremental Sync
- Timing: 04:00 Uhr täglich (nach Borg)
- Paralleler: 4 Transfers, 8 Checkers
8. Hetzner S3
- Bucket:
mbo-backups - Struktur:
production/repos/{hostname}/{archive} - Versionierung: Optional aktivierbar
- Lifecycle: Optional für alte Versionen
Datenfluss
┌──────────────────────────────────────────────────────────────┐
│ KESTRA TIMER │
│ Täglich 02:00 Uhr │
└────────────────┬─────────────────────────────────────────────┘
│
▼
┌──────────────────────────────────────────────────────────────┐
│ NETBOX API QUERY │
│ GET /api/dcim/devices/?tag=backup:true │
│ → [{name, ip, os, status}, ...] │
└────────────────┬─────────────────────────────────────────────┘
│
▼
┌──────────────────────────────────────────────────────────────┐
│ ANSIBLE INVENTORY GENERATION │
│ group_vars, host_vars aus Netbox-Daten │
└────────────────┬─────────────────────────────────────────────┘
│
▼
┌──────────────────────────────────────────────────────────────┐
│ ANSIBLE PLAYBOOK (PARALLEL) │
│ • install-borg-client.yaml │
│ • Per Host: SSH-Key, Scripts, Systemd Timer │
└────────────────┬─────────────────────────────────────────────┘
│
┌─────────┴─────────┐
│ │
▼ ▼
┌──────────────────┐ ┌──────────────────────┐
│ CLIENT HOST 1 │ │ CLIENT HOST 2 │
│ • borg --version│ │ • borg --version │
│ • systemd timer │ │ • systemd timer │
│ • SSH configured│ │ • SSH configured │
└────────┬─────────┘ └──────────┬───────────┘
│ │
│ 03:00 Uhr │
│ (Systemd Timer) │
│ │
▼ ▼
┌────────────────────────────────────┐
│ Borg Backup Execution (LOKAL) │
│ • pre-backup-docker.sh │
│ • MySQL/PostgreSQL Dumps │
│ • Volume Exports │
│ • borg create --stats │
│ • borg prune (retention) │
└────────┬─────────────────────────────┘
│
│ SSH zum Backup-Server
│
▼
┌──────────────────────────────────────────────────────────────┐
│ BACKUP-SERVER │
│ SSH: backup@backup.pve.mbo-tech-it.de │
│ Repos: /backup/repos/{hostname}/ │
│ • Empfängt Borg Archive via SSH │
│ • Speichert in lokalen Repositories │
│ • Pre-checks für Deduplizierung │
└────────────────┬─────────────────────────────────────────────┘
│
│ 04:00 Uhr (Systemd Timer)
│
▼
┌──────────────────────────────────────────────────────────────┐
│ RCLONE SYNC │
│ rclone sync /backup/repos hetzner:mbo-backups/... │
│ • Incremental: nur neue/veränderte Blöcke │
│ • Parallel: 4 transfers │
│ • Log: /var/log/mbo-backup/rclone-sync.log │
└────────────────┬─────────────────────────────────────────────┘
│
▼
┌──────────────────────────────────────────────────────────────┐
│ HETZNER S3 │
│ Bucket: mbo-backups │
│ Struktur: production/repos/{hostname}/{archive-name} │
│ • Redundant in mindestens 2 Datacentern │
│ • Optional: Versioning, Lifecycle Policies │
└──────────────────────────────────────────────────────────────┘
Timing & Ablauf
| Zeit | Komponente | Aktion |
|---|---|---|
| 02:00 | Kestra Timer | Startet Workflow |
| 02:05-02:30 | Netbox API | Lädt Liste aller zu backupenden Hosts |
| 02:30-02:45 | Ansible | Deployed/Updated Borg auf allen Clients |
| 03:00 | Systemd Timer | Startet Backup auf allen Clients |
| 03:00-03:45 | Borg | Backups laufen parallel auf allen Hosts |
| - | Docker Hooks | MySQL/PG Dumps, Volume Exports |
| - | SSH | Archivs gehen an Backup-Server |
| 04:00 | Rclone Timer | Startet S3-Sync |
| 04:00-04:30 | Rclone | Sync zu Hetzner S3 |
Skalierbarkeit
Horizontal (mehr Hosts)
- Borg: Parallelisierung via SSH
- Kestra: Workflow-Engine skaliert linear
- Rclone: Multi-threaded, weitere --transfers falls nötig
- S3: Unbegrenzt
Vertikal (größere Datenmengen)
- Borg: Kompression + Deduplizierung spart 50-70%
- Backup-Server: NVMe für Speed, großes Volume für Speicher
- S3: Skalierung automatic
Parallelität
Max parallel Backups: Anzahl der Clients × CPU-Kerne
Empfehlung: Nice-Priorität senken, CPU-Limits setzen
Sicherheit
Verschlüsselung
- Borg: AES-256-CTR Encryption
- Passphrase: Via Ansible Vault gespeichert
- Transport: SSH (Ed25519 Keys)
Zugriffskontrolle
- SSH Keys: Pro Client eindeutig
- SSH Restrictions:
restrict,command=...auf Server - Backup-User: Minimal privileges
- S3 Credentials: Via Vault/Secrets
Audit & Logging
- Journald: Alle Backup-Logs in Systemd
- Dateilogging:
/var/log/mbo-backup/*.log - Rclone: Detailed Logs für S3-Sync
- Kestra: Workflow Execution History
Fehlerbehandlung
Backup-Fehler
- Borg versucht 3x neu (systemd Restart)
- Fehler geht in Log
- Kestra prüft Exit-Code
- Optional: Notification an Admin
SSH-Fehler
- SSH retry mit exponential backoff
- Key-Rotation wenn nötig
- Fallback: Manual trigger
S3-Fehler
- Rclone retry mit backoff
- Partial sync wird fortgesetzt
- Alert bei persistent failure
Monitoring & Observability
Metriken
- Backup-Größe pro Host
- Deduplizierung-Ratio
- Execution-Zeit
- S3-Upload-Speed
- Retention-Policy Compliance
Logs
- Systemd Journal:
journalctl -u mbo-backup - File-basiert:
/var/log/mbo-backup/*.log - Kestra: Execution History
- Rclone:
--log-level INFO/DEBUG
Alerting
- Backup Failure → Alert
- S3-Sync Failed → Alert
- Disk Space Low → Alert
- SSH Key Expiry → Alert (optional)
Disaster Recovery
RTO/RPO Zielsetzungen
- RTO (Recovery Time Objective): < 2 Stunden
- RPO (Recovery Point Objective): < 24 Stunden
Restore-Optionen
- Lokal vom Backup-Server: Schnell, kein Internet nötig
- Von Hetzner S3: Geo-redundant, asynchron
- Selective Restore: Einzelne Dateien/Volumes
Test-Strategie
- Monatlich: Restore-Drill eines Vollsystems
- Quarterly: S3-Restore-Test
- Nach Major Changes: Schnell-Sanity-Check
Version: 1.0
Stand: 2026-09-09
Nächste Review: 2026-12-09