Dev Tool Experiences
All articles

· 7 min read

Startup Pivots to Self-Hosting LLMs, Discovers It Is Now a GPU Rental Company

By P. Taylor

  • tools
  • satire

SATIRE — On Monday, the startup CopperKite announced that it had achieved model sovereignty. By Tuesday, it had also achieved a ticket queue called “GPU 3 is hot again,” a spreadsheet called “who owns GPU 3,” and a new business unit whose product was waiting for GPU 3 to become emotionally available. The company began the quarter as a workflow software startup. It ended the quarter as a regional appliance of compute allocation, with a small workflow software hobby on the side.

The strategic decision: stop renting intelligence, start renting metal

The original plan was familiar: replace a metered model API with a self-hosted model, keep customer data close, control latency, and turn one variable bill into a predictable operating expense. This was presented in the finance deck as a simple substitution: arrows pointed from “tokens” to “GPU.” The arrows were correct in the same way a map from a kitchen to the ocean is correct. There is, technically, a route.

The first successful local request arrived 11 seconds after docker compose up. Everyone in the incident channel celebrated. The model produced an acceptable summary of a support ticket. The next request arrived 14 seconds later and was rejected because the container had been restarted by someone updating a YAML file that had become, without consultation, the company’s capacity plan.

Docker Compose can reserve GPU devices for a service, including a specific device_ids list or a device count; it also requires capabilities: [gpu]. CopperKite interpreted this as an invitation to write a 137-line compose file called inference-final-final-please.yaml. The file assigned GPU 0 to embeddings, GPU 1 to chat, GPU 2 to a staging environment, and GPU 3 to “experiments,” a category later found to contain one unattended terminal running a 70-billion-parameter model against a README.

services:
  model:
    image: local-inference:latest
    deploy:
      resources:
        reservations:
          devices:
            - driver: nvidia
              device_ids: ["0"]
              capabilities: [gpu]

This was the last moment when the company could honestly say it was “just running a model.”

Every feature request becomes a scheduling policy

Then Sales asked for a demo environment. Support asked for a private deployment. Research asked for an overnight fine-tune. An engineer asked whether they could borrow one card for 20 minutes to test quantization. The answer was yes, provided they first locate the person who owns the card, the person who owns the process using the card, and the person who owns the Slack thread explaining why the process cannot be stopped.

At this point, CopperKite upgraded from Compose to Kubernetes, because the organization needed an industrial-strength way to discover that every useful accelerator was already busy. Kubernetes exposes GPUs through vendor device plugins and schedules them as resources such as nvidia.com/gpu; it also expects operators to install the relevant drivers and device plugin. This converted a deployment concern into a governance framework with YAML indentation.

The platform team introduced node labels: accelerator=large, accelerator=larger, and accelerator=do-not-touch-unless-you-enjoy-a-postmortem. This was not technically required, but it gave the capacity dispute a stable interface. A pod requesting one GPU would remain Pending for 46 minutes, which was considered a major reliability improvement over the prior system, in which it would run immediately on the wrong machine and make the office lights dim.

Observability, or watching money develop a temperature

The new daily stand-up began with the command everyone learned to fear:

nvidia-smi --query-gpu=index,name,utilization.gpu,memory.used,power.draw,temperature.gpu --format=csv,noheader

The command was intended to provide clarity. Instead, it created a new class of questions: Why is memory full when utilization is 0%? Why does the test cluster use more power than production? Is GPU 2 serving traffic, compiling kernels, or grieving? NVIDIA’s own documentation notes that reported frame-buffer memory can be affected by driver reservations, ECC, and operating-system memory accounting, which allowed every graph to become both alarming and deniable.

CopperKite’s dashboard acquired the visual language of an energy utility. It showed memory use, power draw, temperatures, queue depth, model replicas, and a large green rectangle labeled “Cost Avoidance.” The last panel was driven by a shell script that subtracted hypothetical API spending from actual infrastructure spending, then stopped before infrastructure spending could answer back.

The cloud bill returns wearing a hard hat

Self-hosting did reduce one kind of variance. It replaced the surprise of per-request charges with the steadier experience of paying for capacity while usage was low, then paying for more capacity when usage became high. The finance team, which had been promised predictability, received a forecast containing reserved instances, on-demand instances, interrupted spot capacity, persistent volumes, egress, snapshots, observability, and a line item named “temporary temporary node.”

This is where the company discovered its actual product. Not the chat assistant, which was still in beta. Not the workflow automation suite, which had been moved behind a feature flag. The product was a calendar system for expensive rectangles. Engineers were now customers. GPUs were inventory. The platform team was dispatch. The CTO, previously responsible for technical strategy, had become the superintendent of a very quiet rental yard.

The bad part is not that self-hosting is impossible. It is that it is unusually good at revealing work that managed APIs had been doing invisibly: capacity planning, admission control, isolation, rollouts, observability, recovery, and deciding whose request matters when the hardware is full. A local model can be cheaper, faster, more private, or simply the right operational choice. It can also turn a software company into a company whose most urgent customer question is: “Can I have one GPU until lunch?”

The true observation beneath the joke is plain: self-hosting a model is not merely a model decision. Once more than one workload needs the hardware, it is an infrastructure and operations decision—and it should be staffed, measured, and priced that way.

Sources & citations

  1. [1]Docker Docs — Run Docker Compose services with GPU access
  2. [2]Kubernetes Docs — Schedule GPUs
  3. [3]NVIDIA Docs — NVIDIA System Management Interface (nvidia-smi)
  4. [4]Amazon EC2 Pricing