↓ Skip to main content

Project Aether: Building a Production-Grade Homelab on a Repurposed Laptop

Chetan Thapliyal
Author
Chetan Thapliyal
Cloud and DevOps Engineer with a passion for writing.
Table of Contents
homelab - This article is part of a series.
Part : This Article

For a while my home server was a single machine running Docker Compose. It worked fine until it didn’t: one bad container update, and everything on that box went down together. No isolation, no rollback, no way to test changes without touching production. The obvious fix was to restructure it properly, so I did, and that restructuring became Project Aether.

The goal: migrate 34 Docker Compose services to a production-grade Kubernetes platform, with GitOps, encrypted secrets in version control, and proper failure isolation – running entirely on one repurposed HP laptop.

This post is the series opener. It covers what I’m running on, how the architecture fits together, and the reasoning behind the major decisions. The rest of the series goes through each layer in detail.


The hardware
#

The entire platform runs on one machine:

Component Spec
Model HP Notebook (repurposed)
CPU Intel Core i3-5005U, 2 cores / 4 threads, 1.90 GHz
RAM 12 GB DDR3
Storage (OS) 256 GB SSD – system disk, VMs, K3s nodes
Storage (data) 500 GB HDD – ZFS pool aether-pool, media library, backups
Network Gigabit Ethernet + Tailscale VPN
Hypervisor Proxmox VE 9.2, kernel 7.0.14-pve

Display and lid removed. It sits in a Thermaltake case for airflow and runs headless around the clock.

The i3-5005U is a 2015 Broadwell-U chip with 15 W TDP. It has enough for this workload, but there is no headroom to waste. Every architecture decision in this project flows from that constraint.


Why Proxmox
#

The first decision was what to run on bare metal.

The obvious alternative was Ubuntu with Docker directly on the host. That approach works but has a specific failure mode: everything shares one kernel and one root filesystem. A misbehaving container, a bad apt upgrade, or a kernel panic takes down the entire machine including the network stack and remote access layer. There is no recovery without physical access.

Proxmox gives each workload its own VM boundary with its own kernel. A panic inside the K3s VM does not affect anything else. More practically, Proxmox supports instant hypervisor-level snapshots. Before a risky change I take a snapshot, apply the change, and roll it back in seconds if something breaks. That is worth more than the small performance cost of virtualization overhead.

VMware ESXi was considered and dropped. Broadcom’s post-acquisition licensing changes closed the free tier, and the resource footprint on an i3 is heavier than Proxmox’s KVM/LXC hybrid.


Storage: two tiers, not one pool
#

The laptop has two drives with very different characteristics. The decision was how to use them.

The SSD is fast but small. The HDD has four times the capacity but mechanical seek times that make K3s miserable (its embedded SQLite datastore is sensitive to I/O latency). Running everything on one pool would throttle either capacity or performance.

The split:

  • SSD – Proxmox host OS, VM root disks, K3s node disks, container images, database working data
  • HDD – media library (Jellyfin), Proxmox vzdump backups, ZFS pool for bitrot protection on cold data

A ZFS mirror across both drives was considered and rejected. A mirror runs at the speed of the slower drive and is capped at the size of the smaller one. That would waste the HDD’s capacity and destroy the SSD’s performance advantage simultaneously.

The downside of the two-tier approach is that every disk assignment is explicit – Terraform VM specs, PVC storage classes, and backup targets all have to specify which pool. That adds configuration complexity, but the alternative is a cluster that crawls under any I/O load.


K3s over Talos and MicroK8s
#

With Proxmox as the hypervisor, the next question was which Kubernetes distribution to run inside the VMs.

Talos was appealing. It is API-driven, immutable, and purpose-built for Kubernetes. The problem is that it removes SSH access and replaces standard Linux tooling with its own API surface. On a single-operator homelab where the primary goal includes learning to debug Kubernetes, losing the ability to SSH into a node and inspect things with standard tools is a real cost, not a theoretical one.

MicroK8s runs on snap packages and is heavily tied to the Canonical ecosystem. K3s is more common in real-world edge deployments and maps more directly to what production environments at the SMB and edge scale actually use.

K3s won on three counts: it compiles everything into a single binary with a small memory footprint, it runs on a standard Ubuntu VM with full SSH access, and the skills transfer to production edge clusters. The trade-off is that the embedded SQLite datastore gives no control-plane high availability – if the control plane VM dies, the API is unreachable until it comes back. For a one-person homelab that is acceptable.

The cluster is three nodes: one control plane (2 vCPU, 2 GB), two workers (2 vCPU, 3.5 GB each).


Tailscale for network access
#

The homelab is behind a residential ISP in India. Most residential connections here use Carrier-Grade NAT, which means there is no public IP address to forward ports to, even if opening ports to the internet were a good idea. It is not.

The requirement was remote access from anywhere – SSH to nodes, the Proxmox web UI, the K3s API – without exposing anything publicly.

Cloudflare Tunnels handle inbound traffic well but are designed for publishing services to the internet, not for acting as a management VPN. Port forwarding was ruled out immediately.

Tailscale uses WireGuard under the hood and handles NAT traversal automatically, including CGNAT. Every node and operator device joins the same mesh network. SSH access, the Proxmox UI, and the K3s API endpoint all become reachable over their Tailscale MagicDNS names with no open ports on the router. The only dependency is the Tailscale coordination plane, which continues passing peer-to-peer traffic even if it goes down – new nodes just cannot join while it is unavailable.

The auth key is encrypted with SOPS before it touches the Git repository. More on that in a later post.


The full stack
#

Layer Tool Purpose
Hypervisor Proxmox VE 9.2 Bare-metal virtualisation, isolation, snapshots
IaC Terraform + bpg/proxmox VM provisioning from Cloud-Init template
State backend GCS (tf-backend-oci-hlab-gcs) Remote Terraform state
Configuration Ansible K3s install and cluster bootstrap
Orchestration K3s v1.36 Lightweight Kubernetes
GitOps ArgoCD Declarative app delivery from Git
Ingress Ingress-NGINX + MetalLB Load balancing and routing
TLS cert-manager + Cloudflare Automatic HTTPS
Secrets SOPS + age Encrypted secrets in Git
Networking Tailscale Remote access and inter-node mesh
Automation n8n Workflow automation
Task runner Task One command for every operation
CI GitHub Actions Pre-commit, lint

What is being migrated
#

The existing setup is 34 Docker Compose services, currently archived in legacy-docker/. They cover media (Jellyfin, Navidrome, Audiobookshelf), productivity (Siyuan, Karakeep, Mealie, Stirling-PDF), and a few infrastructure pieces (Portainer, tsdproxy) that will be replaced by their Kubernetes equivalents.

Not everything is moving. Portainer is replaced by ArgoCD. tsdproxy is replaced by the Tailscale Operator. Nextcloud, Immich, and Plex are deferred – they are either too resource-heavy or need rethinking on limited RAM. The realistic target is 25 services migrated across five waves, ordered by complexity.


Where Phase 1 ended
#

Phase 1 is complete. The cluster is running:

  • Proxmox VE 9.2 installed on the SSD, ZFS pool on the HDD
  • Cloud-Init template (VM 9000) built from Ubuntu 24.04
  • Three K3s nodes provisioned via Terraform, static IPs, GCS remote state
  • Cluster bootstrapped via Ansible, Tailscale installed on all nodes
  • SOPS + age configured, secrets encrypted before they touch Git
  • Taskfile and pre-commit hooks in place across the repository

Each of those steps has its own post in this series. Phase 2 is the GitOps layer: ArgoCD, the App-of-Apps pattern, MetalLB, and cert-manager.


Architecture
#

graph TD
    DEV["Developer (Pulsar)"]

    subgraph HOMELAB["Homelab — Proxmox VE 9.2"]
        PVE["Proxmox Host 192.168.1.200"]

        subgraph K3S["K3s Cluster"]
            CTL["k3s-control-01 192.168.1.51 2 vCPU / 2 GB"]
            W1["k3s-worker-01 192.168.1.52 2 vCPU / 3.5 GB"]
            W2["k3s-worker-02 192.168.1.53 2 vCPU / 3.5 GB"]
        end

        subgraph GITOPS["GitOps Layer"]
            ARGO["ArgoCD"]
            APPS["Apps: n8n, Gitea, Vaultwarden, Immich, more"]
        end

        subgraph SYSTEM["Platform Services"]
            LB["MetalLB"]
            ING["Ingress-NGINX"]
            CERT["cert-manager"]
            MON["Prometheus + Grafana"]
        end
    end

    subgraph NETWORK["Network"]
        TS["Tailscale Mesh VPN"]
        CF["Cloudflare DNS + Tunnels"]
    end

    DEV -- "terraform apply / ansible-playbook / kubectl" --> PVE
    DEV -- "Tailscale" --> TS
    TS -- "SSH / API" --> PVE
    PVE -- "Cloud-Init clone" --> CTL & W1 & W2
    CTL -- "manages" --> W1 & W2
    ARGO -- "sync from Git" --> APPS
    ARGO --> SYSTEM
    CF -- "DNS + TLS" --> ING
    LB --> ING
homelab - This article is part of a series.
Part : This Article

Related