Skip to content
All work

MLOps, Infrastructure & Tooling

On-Prem AI Server Infrastructure

HawkEyes · internal infrastructure + vastai-automate · 2025–present

Two self-hosted Ubuntu GPU servers, designed and built in-house — an alternative to renting cloud GPUs.

Ubuntu ServerLinux (Arch)NVIDIA GPUDockerBashVast.ai
2
on-prem GPU servers(CV)
3+ yrs
Arch Linux daily driver(CV)
in-house
vs rented cloud GPU(CV)

Problem

Relying on rented cloud GPUs for hosting AI services and training models is costly and adds an external dependency the team can't fully control.

Approach

  • 01Designed, built, and administer two dedicated on-premises Ubuntu servers — hardware provisioning, Linux server setup, containerized deployment, and ongoing maintenance.
  • 02Host most of the team's production AI workloads and run model-training jobs in-house.
  • 03Complement on-prem capacity with vastai-automate: scripts that search Vast.ai offers under hardware/network/reliability constraints, rank by price/throughput/reliability, auto-select the best instances, and bootstrap the environment.

Results & Impact

  • Reduced dependence on rented cloud GPUs and gave the team full control over training and inference infrastructure.— CV-sourced
  • Built and maintained by a long-time Linux power user (Arch daily driver 3+ years; hands-on across Ubuntu, Fedora, Mint, Pop!_OS).— CV-sourced