Building a Self-Hosted GitHub Mirror and Disaster-Recovery Backup with Forgejo, Kopia, Docker Desktop and Windows
The first piece I wanted to solve was source code.
My GitHub account contains hundreds of repositories accumulated over many years. GitHub is still the primary working copy because it integrates with my development workflow, CI/CD and tools, but I wanted:
- a complete local mirror of every repository I own;
- automatic discovery of newly created GitHub repositories;
- encrypted versioned backups;
- protection against accidental deletion or corruption;
- a recovery process I had actually tested;
- something another technical person could rebuild from documentation.
The resulting architecture is:
GitHub │ │ Pull mirrors ▼ Forgejo │ │ Consistent local archive ▼ Kopia │ │ AES-256 encrypted / deduplicated ▼ Large local backup disk
The important part is that I did not stop when the backup command said success.
I actually restored the backup into a new Docker volume, launched a completely separate Forgejo server from it, logged in, and verified that all repositories were present and usable.
This article documents the entire process.
The Goal
My starting point was approximately:
GitHub repositories: 358 Total Git repository data: ~6.2 GB Windows system disk free: ~333 GB Backup disk: 16 TB Backup disk free: ~9 TB
I wanted GitHub to remain the development system of record:
Developer machines │ ▼ GitHub │ ▼ Forgejo
Rather than changing every development machine to push into Forgejo.
This has several advantages:
- existing GitHub workflows remain unchanged;
- GitHub Actions continue to work;
- Codex and other GitHub integrations continue to work;
- Forgejo becomes a passive local preservation layer;
- recovery is independent of GitHub being available.
Why Forgejo?
Forgejo is a lightweight self-hosted Git service descended from Gitea.
For this project it gives me:
- Git repository hosting;
- repository mirroring;
- private local copies of public and private repositories;
- SQLite support;
- a simple Docker deployment;
- low resource requirements;
- a familiar GitHub-like interface.
Most importantly, Forgejo supports pull mirrors.
A pull mirror periodically synchronizes from GitHub:
GitHub repository │ │ git fetch ▼ Forgejo mirror
This is better than periodically running arbitrary clone scripts because Forgejo understands that the repository is a mirror and manages synchronization itself.
Important distinction: replication is not backup
This is worth stating explicitly.
A mirror is not a backup.
Suppose this happens:
Monday GitHub repo exists │ ▼ Forgejo mirror exists
Then on Tuesday:
GitHub repo accidentally deleted
A synchronization system exists to reproduce upstream state.
What I really wanted was:
GitHub │ ▼ Forgejo │ ▼ Kopia snapshots │ ├── yesterday ├── last week ├── last month └── last year
That is why Forgejo and Kopia solve two different problems.
Part 1 — Installing Forgejo
My home server runs Windows with Docker Desktop and WSL2.
I created:
mkdir C:\Homelab mkdir C:\Homelab\forgejo cd C:\Homelab\forgejo
Before assigning ports, I checked they were unused:
Get-NetTCPConnection -LocalPort 3000,2222 -State Listen -ErrorAction SilentlyContinue
Then I created:
C:\Homelab\forgejo\compose.yaml
with:
services: forgejo: image: codeberg.org/forgejo/forgejo:15.0.8 container_name: forgejo environment: USER_UID: "1000" USER_GID: "1000" FORGEJO__database__DB_TYPE: sqlite3 FORGEJO__server__SSH_PORT: "2222" restart: unless-stopped volumes: - forgejo-data:/data ports: - "3000:3000" - "2222:22" volumes: forgejo-data: name: forgejo-data
Start it:
docker compose pull docker compose up -d
Check it:
docker compose ps docker logs forgejo --tail 50
Forgejo becomes available at:
http://localhost:3000
Why use a Docker named volume?
Forgejo stores its complete application state under:
/data
That includes much more than Git repositories:
/data ├── git/ │ └── repositories/ ├── gitea/ │ ├── gitea.db │ ├── conf/ │ ├── attachments/ │ ├── avatars/ │ ├── packages/ │ └── ... └── ssh/
Using:
forgejo-data:/data
keeps this inside Docker's Linux filesystem.
That is particularly useful for SQLite and application metadata, rather than putting a live database directly onto an NTFS or exFAT bind mount.
Part 2 — Creating the GitHub read-only token
I created a GitHub fine-grained Personal Access Token.
The important principle was least privilege.
I gave it:
Repository access: All repositories owned by my account Permissions: Contents: Read-only Metadata: Read-only
Nothing else was required for normal Git cloning.
There is no reason for this mirror token to be able to:
write code delete repositories manage Actions manage secrets modify webhooks administer repositories
Part 3 — Testing one repository manually
Before automating hundreds of repositories, I tested one.
In Forgejo:
New Migration → Git → GitHub repository URL → Authentication token → This repository will be a mirror → Make repository private → Migrate
After migration I checked:
Repository → Settings → Repository → Mirror Settings
The critical result was:
Direction: Pull Synchronize now
That proves it is an actual Forgejo pull mirror.
A normal repository migration is not the same thing.
Importantly, Forgejo does not simply turn an existing normal repository into a pull mirror later. The repository should be created as a mirror from the beginning.
Part 4 — Automatically mirroring every GitHub repository
Doing this manually 358 times would obviously be unreasonable.
I created:
C:\Homelab\github-forgejo-sync
containing:
compose.yaml sync.py .env .gitignore
Environment configuration
My .env looks conceptually like this:
GITHUB_OWNER=YOUR_GITHUB_USERNAME GITHUB_TOKEN=YOUR_READ_ONLY_GITHUB_TOKEN FORGEJO_URL=http://host.docker.internal:3000 FORGEJO_OWNER=your-forgejo-user FORGEJO_TOKEN=YOUR_FORGEJO_API_TOKEN FORGEJO_FORCE_PRIVATE=true MIN_FREE_GB=100 MAX_NEW_REPOS=0 DRY_RUN=true
Never commit .env.
My .gitignore therefore contains:
.env
Why host.docker.internal?
Inside a Docker container:
localhost
means the container itself.
It does not mean the Windows machine.
Docker Desktop provides:
host.docker.internal
for reaching the host.
So the Forgejo API URL from the synchronization container becomes:
http://host.docker.internal:3000
Safety rules in the importer
I deliberately made the synchronization script conservative.
It:
enumerates GitHub repositories │ ▼ filters owner == my GitHub username │ ▼ checks Forgejo │ ├── already exists → skip │ └── missing → create pull mirror
I also added:
MIN_FREE_GB=100
so imports stop if the system disk becomes unexpectedly full.
Another important safeguard was:
repo["owner"]["login"].lower() == GITHUB_OWNER.lower()
This prevents repositories merely shared with me—employer repositories, organizations, collaborations, etc.—from silently being mirrored.
Dry run first
Before creating anything:
DRY_RUN=true
Then:
docker compose run --rm github-forgejo-sync
My result was:
GitHub repos : 358 Existing Forgejo mirrors : 1 New mirrors created : 0 Existing non-mirrors : 0 Failures : 0
The script also calculated:
GitHub reported repo size : ~6,242 MB
This was reassuring because I had over 300 GB free on C:.
First controlled real import
Rather than immediately importing everything:
DRY_RUN=false MAX_NEW_REPOS=5
Run again:
docker compose run --rm github-forgejo-sync
Then check storage:
docker exec forgejo du -sh /data docker exec forgejo du -sh /data/git/repositories
Initial state:
/data 12.9 MB repositories 9.3 MB
After five repositories:
/data 185 MB repositories 178 MB
Everything behaved normally.
Importing everything
I then changed:
MAX_NEW_REPOS=0
where zero means unlimited.
The final result was:
GitHub repos : 358 Existing Forgejo mirrors : 6 New mirrors created : 352 Existing non-mirrors : 0 Failures : 0
The complete Forgejo data volume ended up at:
/data 6.1 GB repositories 6.1 GB
That was remarkably close to GitHub's own reported repository size.
Part 5 — Adding real backups with Kopia
Now the important part.
Forgejo protects me against GitHub availability problems.
Kopia protects me against:
- accidental deletion;
- local corruption;
- bad synchronization;
- database damage;
- Docker volume loss;
- operator mistakes;
- historical recovery.
Kopia provides:
encryption deduplication snapshot history retention policies integrity verification restore
Installing Kopia in Docker
I created:
C:\Homelab\kopia
and:
mkdir C:\Homelab\kopia\config mkdir C:\Homelab\kopia\cache mkdir C:\Homelab\kopia\logs mkdir C:\Homelab\kopia\restore mkdir C:\Homelab\kopia\staging
The Kopia image used was:
kopia/kopia:0.23.1
Check:
docker compose run --rm kopia --version
which returned:
0.23.1
The interesting problem: Docker Desktop would not mount the backup drive
This became the most useful troubleshooting lesson in the entire project.
My large NTFS backup disk was:
E: ~16 TB
I originally tried:
- E:/FamilyBackup/Kopia/HomeServer:/repository
Kopia reported:
Storage capacity: 132 MB Storage available: 56 MB
Clearly wrong.
Inside Docker:
docker compose run --rm --entrypoint sh kopia ` -c "df -h /repository"
returned approximately:
/dev/sdd 127M
while Windows showed a multi-terabyte filesystem.
Even this test:
docker run --rm ` --mount type=bind,source="E:\",target=/mnt/e ` alpine ` sh -c "df -h /mnt/e"
still exposed only the tiny filesystem.
A 256 MB write failed after approximately 60 MB:
No space left on device
So this was definitely not just incorrect free-space reporting.
Docker Desktop was simply not mounting the real E: filesystem.
The workaround: SMB/CIFS
Instead of mounting E: directly:
Windows E: │ ▼ Docker bind mount
I exposed the backup location through Windows SMB:
E:\FamilyBackup\Kopia\HomeServer │ ▼ Windows SMB │ ▼ Docker CIFS volume │ ▼ Kopia
This completely bypassed the broken Docker Desktop drive translation.
Testing SMB first
I created:
New-Item ` -ItemType Directory ` -Path "E:\FamilyBackup\DockerMountTest" ` -Force
Then shared it:
$account = "$env:USERDOMAIN\$env:USERNAME" New-SmbShare ` -Name "DockerMountTest" ` -Path "E:\FamilyBackup\DockerMountTest" ` -FullAccess $account
Created a marker:
"REAL E DRIVE" | Set-Content E:\FamilyBackup\DockerMountTest\marker.txt
My host LAN address was used for SMB access.
Use your own:
Get-NetIPAddress -AddressFamily IPv4 | Where-Object { $_.IPAddress -notlike "127.*" -and $_.IPAddress -notlike "169.254.*" } | Select-Object InterfaceAlias,IPAddress
Create a temporary CIFS Docker volume
Conceptually:
docker volume create ` --driver local ` --opt type=cifs ` --opt "device=//HOST_IP/DockerMountTest" ` --opt "o=addr=HOST_IP,username=USERNAME,password=PASSWORD,vers=3.0,rw" ` smb-test
Then:
docker run --rm ` -v smb-test:/mnt/test ` alpine ` sh -c "df -h /mnt/test; ls -la /mnt/test; cat /mnt/test/marker.txt"
This time Docker reported:
~15 TB total ~8 TB available
and:
REAL E DRIVE
Success.
I also performed a 512 MB write:
docker run --rm ` -v smb-test:/mnt/test ` alpine ` sh -c "dd if=/dev/zero of=/mnt/test/docker-test.bin bs=1M count=512 && sync && ls -lh /mnt/test/docker-test.bin && rm /mnt/test/docker-test.bin"
That succeeded too.
Do not use your normal Windows account permanently
For testing I used my normal account.
For permanent operation I created a dedicated Windows account:
kopia_backup
with access only to the Kopia repository folder.
Create it from elevated PowerShell:
$KopiaPassword = Read-Host ` "Password for kopia_backup" ` -AsSecureString New-LocalUser ` -Name "kopia_backup" ` -Password $KopiaPassword ` -Description "Docker Kopia backup repository" ` -PasswordNeverExpires ` -UserMayNotChangePassword
Give it access:
icacls "E:\FamilyBackup\Kopia\HomeServer" ` /grant "${env:COMPUTERNAME}\kopia_backup:(OI)(CI)M"
Then share only the repository:
New-SmbShare ` -Name "KopiaHomeServer" ` -Path "E:\FamilyBackup\Kopia\HomeServer" ` -FullAccess "${env:COMPUTERNAME}\kopia_backup"
Create the permanent Docker CIFS volume
Conceptually:
docker volume create ` --driver local ` --opt type=cifs ` --opt "device=//HOST_IP/KopiaHomeServer" ` --opt "o=addr=HOST_IP,username=kopia_backup,password=PASSWORD,vers=3.0,rw" ` kopia-repository
Important security note:
Docker stores volume options.
Therefore do not use your main Windows account here.
Use a dedicated account with minimal access.
Final Kopia Compose configuration
The important parts of my compose.yaml are:
services: kopia: image: kopia/kopia:0.23.1 hostname: kopia-homeserver env_file: - .env volumes: - ./config:/app/config - ./cache:/app/cache - ./logs:/app/logs - kopia-repository:/repository - forgejo-data:/source/forgejo:ro - C:/Homelab/forgejo:/source/forgejo-config:ro - C:/Homelab/kopia/staging:/source/staging:ro - ./restore:/restore restart: "no" volumes: forgejo-data: external: true kopia-repository: external: true
Stable Kopia hostname matters
This line is important:
hostname: kopia-homeserver
Otherwise:
docker compose run
creates random container hostnames.
Kopia associates snapshots and maintenance ownership with:
user@hostname
A stable hostname therefore gives:
root@kopia-homeserver
rather than:
root@random-container-id
Creating the Kopia repository
My .env contains:
KOPIA_PASSWORD=A_LONG_RANDOM_SECRET
Do not put the real password in source control.
Then:
docker compose run --rm kopia ` repository create filesystem ` --path=/repository
Status:
docker compose run --rm kopia repository status
My working setup reported:
Storage capacity: 16 TB Storage available: 8.9 TB Encryption: AES256-GCM-HMAC-SHA256 Hostname: kopia-homeserver
Validate the repository provider
This is worth doing:
docker compose run --rm kopia repository validate-provider
My final result:
Opening equivalent storage connections... Validating storage capacity... Validating writes... Validating reads... Validating metadata... Running concurrency test... All good.
That was the point at which I knew SMB/CIFS was behaving correctly enough for Kopia.
Configure retention
I used:
docker compose run --rm kopia policy set --global ` --keep-latest=10 ` --keep-hourly=24 ` --keep-daily=30 ` --keep-weekly=12 ` --keep-monthly=24 ` --keep-annual=10
That gives me:
Latest: 10 Hourly: 24 Daily: 30 Weekly: 12 Monthly: 24 Annual: 10
Then:
docker compose run --rm kopia maintenance set --owner=me
which establishes:
root@kopia-homeserver
as the maintenance owner.
Why I do not snapshot live SQLite directly
Forgejo uses SQLite in my setup.
A filesystem-level backup while the database is actively changing risks capturing inconsistent state.
Instead I use:
Forgejo running │ ▼ stop Forgejo │ ▼ create filesystem TAR │ ▼ restart Forgejo immediately │ ▼ Kopia backs up the TAR
This gives a clean application-consistent snapshot.
The outage only lasts while the local archive is created.
Kopia can then spend as long as necessary copying the archive to the backup disk while Forgejo is already back online.
Why TAR?
I deliberately preserve /data as a TAR file.
That preserves Linux filesystem metadata far better than copying everything into a normal Windows directory.
It also gives me one self-contained recovery artifact:
forgejo-data.tar
containing:
SQLite database Forgejo configuration Git repositories SSH files attachments avatars LFS data packages application state
Backup script
The essential process is:
$ErrorActionPreference = "Stop" $ForgejoContainer = "forgejo" $StagingDir = "C:\Homelab\kopia\staging" $TarFile = Join-Path $StagingDir "forgejo-data.tar" $TempTarFile = Join-Path $StagingDir "forgejo-data.tar.tmp" $KopiaDir = "C:\Homelab\kopia" Write-Host "==============================================" Write-Host "Forgejo consistent backup" Write-Host "Started: $(Get-Date)" Write-Host "==============================================" $wasRunning = $false try { $running = docker inspect ` --format='{{.State.Running}}' ` $ForgejoContainer 2>$null $wasRunning = ($running -eq "true") if (-not $wasRunning) { throw "Forgejo is not currently running." } Write-Host "Stopping Forgejo..." docker stop $ForgejoContainer if ($LASTEXITCODE -ne 0) { throw "Failed to stop Forgejo." } Remove-Item ` $TempTarFile ` -Force ` -ErrorAction SilentlyContinue Write-Host "Creating consistent Forgejo archive..." docker run --rm ` -v "forgejo-data:/source:ro" ` -v "C:/Homelab/kopia/staging:/staging" ` alpine ` sh -c "tar -C /source -cf /staging/forgejo-data.tar.tmp ." if ($LASTEXITCODE -ne 0) { throw "Failed to create Forgejo archive." } Move-Item ` $TempTarFile ` $TarFile ` -Force } finally { if ($wasRunning) { Write-Host "Restarting Forgejo..." docker start $ForgejoContainer | Out-Null } } Set-Location $KopiaDir Write-Host "Snapshotting Forgejo data..." docker compose run --rm kopia ` snapshot create /source/staging ` --description="Forgejo consistent volume archive" if ($LASTEXITCODE -ne 0) { throw "Kopia data snapshot failed." } Write-Host "Snapshotting Docker configuration..." docker compose run --rm kopia ` snapshot create /source/forgejo-config ` --description="Forgejo Docker configuration" if ($LASTEXITCODE -ne 0) { throw "Kopia configuration snapshot failed." } Write-Host "Removing staging archive..." Remove-Item ` $TarFile ` -Force ` -ErrorAction SilentlyContinue Write-Host "Backup completed successfully."
The cleanup occurs only after both snapshots succeed.
That is intentional.
If the backup fails, the staging TAR remains available for investigation.
The first real backup
The first backup produced approximately:
forgejo-data.tar 6.06 GB
Kopia uploaded approximately:
6.5 GB logical data
in roughly a minute and a half on my system.
It created two snapshots:
/source/staging /source/forgejo-config
The important thing is not the exact speed.
It is that:
Forgejo stopped archive created Forgejo restarted snapshot completed
with no errors.
Does Kopia store another 6 GB every day?
No.
Kopia uses content-addressed deduplication.
Conceptually:
Day 1 6 GB ████████████████████ Day 2 mostly unchanged ████████████████████ ▲ only changed chunks stored Day 3 ████████████████████ ▲ changed chunks only
Each snapshot logically represents the complete state.
Physically, unchanged chunks are reused.
So this is not:
6 GB × 365
unless all 6 GB somehow change every single day.
Because the source is a TAR, deduplication may not be absolutely optimal compared with snapshotting millions of individual files directly, but Kopia's content-defined chunking still works well.
Given a multi-terabyte backup disk, this is a very reasonable tradeoff for a simple, highly recoverable archive.
How to confirm backups exist
List snapshots:
docker compose run --rm kopia snapshot list --all
Typical result:
root@kopia-homeserver:/source/forgejo-config <timestamp> ... root@kopia-homeserver:/source/staging <timestamp> ...
Check repository status:
docker compose run --rm kopia repository status
But a backup is not proven until it restores
This was the most important part of the project.
I did not consider this complete after:
Backup completed successfully
I verified the entire recovery chain.
Step 1 — verify every byte
I ran:
docker compose run --rm kopia snapshot verify ` SNAPSHOT_ID ` --verify-files-percent=100
My output confirmed Kopia read the entire ~6.5 GB snapshot successfully.
Step 2 — restore the archive
I restored the snapshot into:
C:\Homelab\kopia\restore\forgejo-test
using:
docker compose run --rm kopia snapshot restore ` SNAPSHOT_ROOT_ID ` /restore/forgejo-test
The result:
Restored: 1 file 1 directory ~6.5 GB
Then:
Get-Item ` "C:\Homelab\kopia\restore\forgejo-test\forgejo-data.tar"
showed:
~6.06 GB
Step 3 — inspect the archive
Before attempting to boot anything:
docker run --rm ` -v "C:/Homelab/kopia/restore/forgejo-test:/restore:ro" ` alpine ` sh -c "tar -tf /restore/forgejo-data.tar > /tmp/list && grep -E '\.db$' /tmp/list && grep -E 'app\.ini$' /tmp/list && grep -E 'git/repositories' /tmp/list | head"
I confirmed:
./gitea/gitea.db ./gitea/conf/app.ini ./git/repositories/
Step 4 — restore into a new Docker volume
Create a clean test volume:
docker volume create forgejo-restore-test-data
Extract the backup:
docker run --rm ` -v "C:/Homelab/kopia/restore/forgejo-test:/restore:ro" ` -v "forgejo-restore-test-data:/data" ` alpine ` sh -c "cd /data && tar -xf /restore/forgejo-data.tar"
Step 5 — count repositories
I checked:
docker run --rm ` -v "forgejo-restore-test-data:/data:ro" ` alpine ` sh -c "find /data/git/repositories \ -mindepth 2 \ -maxdepth 2 \ -type d \ -name '*.git' | wc -l"
Result:
358
Exactly what I expected.
Step 6 — boot the restored server
Now for the real test.
I launched a second Forgejo instance:
docker run -d ` --name forgejo-restore-test ` -p 127.0.0.1:3001:3000 ` -e USER_UID=1000 ` -e USER_GID=1000 ` -e FORGEJO__server__ROOT_URL=http://localhost:3001/ ` -v "forgejo-restore-test-data:/data" ` codeberg.org/forgejo/forgejo:15.0.8
Then:
docker logs forgejo-restore-test --tail 50
The restored server reported:
SQLite3 support is enabled PING DATABASE sqlite3 ORM engine initialization successful Listen: http://0.0.0.0:3000
I opened:
http://localhost:3001
and there it was.
A completely restored Forgejo instance.
All repositories were present.
I could browse commits and files.
The restore was real.
Cleaning up the recovery test
Once verified:
docker rm -f forgejo-restore-test docker volume rm forgejo-restore-test-data
Remove temporary restored files:
Remove-Item ` "C:\Homelab\kopia\restore\forgejo-test" ` -Recurse ` -Force
Scheduling nightly backups
I chose 03:00.
schtasks /Create ` /TN "Homelab - Forgejo Kopia Backup" ` /SC DAILY ` /ST 03:00 ` /TR "powershell.exe -NoProfile -ExecutionPolicy Bypass -File C:\Homelab\kopia\backup-forgejo.ps1" ` /RL HIGHEST ` /F
Check it:
schtasks /Query ` /TN "Homelab - Forgejo Kopia Backup" ` /V ` /FO LIST
A successful scheduled execution should ultimately report:
LastTaskResult : 0
You can force a test run:
schtasks /Run ` /TN "Homelab - Forgejo Kopia Backup"
Then:
Get-ScheduledTask ` -TaskName "Homelab - Forgejo Kopia Backup" | Get-ScheduledTaskInfo
How I now determine whether the system is healthy
My normal health check is simply:
Windows Scheduler ↓ LastTaskResult = 0 Kopia ↓ recent snapshot timestamp Forgejo ↓ container running
Commands:
Get-ScheduledTask ` -TaskName "Homelab - Forgejo Kopia Backup" | Get-ScheduledTaskInfo
docker compose run --rm kopia snapshot list --all
docker ps --filter "name=forgejo"
Periodically I can additionally run:
docker compose run --rm kopia snapshot verify ` SNAPSHOT_ID ` --verify-files-percent=100
And perhaps once or twice per year repeat the complete disaster-recovery test.
Security lessons
There are several things I would strongly recommend.
Never commit .env
Use:
.env
Secrets include:
GitHub token Forgejo API token Kopia encryption password SMB credentials
Be careful with docker compose config
docker compose config expands environment variables.
That means secrets may appear directly in terminal output.
If terminal output containing a secret is copied somewhere public, rotate that secret immediately.
This is easy to overlook.
Use dedicated service accounts
Do not give Docker your normal Windows login if it only needs backup access.
Create:
kopia_backup
and restrict it to:
E:\FamilyBackup\Kopia\HomeServer
Keep mirrors private
Even if a GitHub repository is public today, I make all local Forgejo mirrors private.
The local estate is not intended to become another public Git hosting service.
What this protects against
This architecture protects reasonably well against:
GitHub outage GitHub account loss accidental repository deletion bad repository changes Docker volume loss Forgejo corruption local system disk failure historical recovery requirements
But one threat remains.
This is not yet full 3-2-1 backup
The backup disk is physically connected to the same computer.
Therefore this scenario is still dangerous:
fire theft electrical disaster ransomware affecting all attached storage physical destruction of the machine and drives
True 3-2-1 protection requires:
3 copies of important data 2 different storage types 1 copy off-site
My eventual design will look like:
GitHub │ ▼ Forgejo │ ▼ Kopia / \ / \ ▼ ▼ local backup encrypted off-site 16 TB
The off-site copy could eventually be:
another physical location Backblaze B2 S3-compatible storage another trusted family server encrypted cloud object storage
What about GitHub metadata?
A Git mirror preserves Git history:
commits branches tags repository files
But GitHub contains additional state that is not inherently part of Git:
Issues Pull Requests Discussions release metadata release assets Actions artifacts GitHub Packages repository settings branch protections secrets
Those require separate archival strategies if they matter.
For my initial goal, preserving source code and complete repository history was the priority.
Git LFS caveat
Git LFS deserves special consideration.
A normal Git repository can contain LFS pointer files while the actual large objects live elsewhere.
If you use Git LFS extensively, explicitly confirm that LFS objects are also being preserved.
Do not assume seeing a .gitattributes file means the binary content itself has been safely backed up.
What I would do differently from the beginning
If I were rebuilding this tomorrow, my order would be:
1. Install Forgejo 2. Create one manual pull mirror 3. Verify pull synchronization 4. Build GitHub discovery/import automation 5. Run dry-run 6. Import five repositories 7. Measure disk usage 8. Import everything 9. Install Kopia 10. Validate backup storage independently 11. Configure retention 12. Build consistent Forgejo archive 13. Run first backup 14. Verify entire snapshot 15. Restore into temporary directory 16. Restore into Docker volume 17. Boot separate Forgejo instance 18. Confirm repositories 19. Only then schedule automatic backups
That progression minimizes risk.
Final result
The finished system looks like this:
DEVELOPMENT Developer machines │ ▼ GitHub │ │ read-only pull mirrors ▼ Forgejo 358 repositories │ │ consistent archive ▼ Kopia │ │ encrypted │ deduplicated │ versioned ▼ 16 TB backup disk
And the disaster-recovery path has been tested:
Kopia snapshot │ ▼ restore archive │ ▼ new Docker volume │ ▼ new Forgejo container │ ▼ SQLite starts │ ▼ 358 repositories found │ ▼ repositories browsable
That final test changed my confidence in the system completely.
There is an enormous difference between:
“I think my backups work.”
and:
“I deleted nothing from production, restored the backup into an empty environment, booted the application from it, and used it.”
The latter is the standard I want for the rest of the family digital estate as well.
Where this goes next
Forgejo was only the first layer.
The wider system I am building will eventually look something like:
FAMILY DIGITAL ESTATE Homepage │ ┌──────────────────┼───────────────────┐ │ │ │ ▼ ▼ ▼ Immich Paperless BookStack photos/video documents family handbook │ │ │ └──────────────────┼───────────────────┘ │ ▼ Kopia │ encrypted snapshots │ ┌──────────────┴──────────────┐ ▼ ▼ local storage off-site
The next component will be Immich, so the same principles can be applied to decades of family photos and videos:
preserve originals do not lock data into one application keep databases separately recoverable back everything up test restoration document the process
That is ultimately the point of the project.
Not just self-hosting applications.
Building a digital estate that can survive hardware failure, cloud-provider changes, mistakes—and eventually even the person who originally built it.

Comments
Post a Comment