MLOps, Infrastructure & Tooling
On-Prem AI Server Infrastructure
HawkEyes · internal infrastructure + vastai-automate · 2025–present
Two self-hosted Ubuntu GPU servers, designed and built in-house — an alternative to renting cloud GPUs.
Ubuntu ServerLinux (Arch)NVIDIA GPUDockerBashVast.ai
- 2
- on-prem GPU servers(CV)
- 3+ yrs
- Arch Linux daily driver(CV)
- in-house
- vs rented cloud GPU(CV)
Problem
Relying on rented cloud GPUs for hosting AI services and training models is costly and adds an external dependency the team can't fully control.
Approach
- 01Designed, built, and administer two dedicated on-premises Ubuntu servers — hardware provisioning, Linux server setup, containerized deployment, and ongoing maintenance.
- 02Host most of the team's production AI workloads and run model-training jobs in-house.
- 03Complement on-prem capacity with vastai-automate: scripts that search Vast.ai offers under hardware/network/reliability constraints, rank by price/throughput/reliability, auto-select the best instances, and bootstrap the environment.
Results & Impact
- Reduced dependence on rented cloud GPUs and gave the team full control over training and inference infrastructure.— CV-sourced
- Built and maintained by a long-time Linux power user (Arch daily driver 3+ years; hands-on across Ubuntu, Fedora, Mint, Pop!_OS).— CV-sourced
Related repositories
vastai-automate
GPU instance selection & provisioning automation (Vast.ai cost optimization).
DockerLessons
Docker & containerization practice (Compose, cheat sheet).
important_files
ML/CV engineering toolbox — dataset prep, YOLO annotation, training utils.
custom-fastfetch
Fastfetch config presets & ASCII art (Arch / Hyprland dotfiles).