ComputerWeekly

The GPU mirage on a lazy Sunday afternoon


I often need to transcribe audio to turn interviews, voice notes etc into plain text. And I am naturally quite loath to spend money on things I can, in theory, do myself. So for a while I had thought about how I could do it without paying a monthly subscription for the privilege. 

The obvious solution is Whisper, OpenAI’s open-source speech-to-text model – effectively a dedicated large language model for transcription, and free to download. Stick it in the cloud, run it on an Ubuntu box, drive it from a Python script, hook it up to a GPU, switch it on when there is work and off when there is not. It’ll cost a few tens of cents a month, perhaps. I have a Google Cloud account, and it is a Sunday afternoon. How hard can it be?

Off I go to my Google Cloud Console and the first frustrations commence. Google gives a new account a GPU quota of zero, so nothing can start until I ask for permission to spend my own money. I ask. The request clears. The requested GPU instance sits behind a spinning blue circle that does not stop. A red status icon announces the zone has exhausted its supply of graphics cards.

So I delete the instance and build it again somewhere else. Then somewhere else. Same result. Virginia, Iowa, London, I try them all. No joy. Then a cheaper, older card instead. Each time the same blue circle, the same red icon. Chatting to Gemini at the same time – yes, I know, talking to a guessing machine – “I’m really surprised there isn’t some tool on here that just allows you to define where the nearest L4 is.”

There is no map to show where the GPUs are. You guess a zone and you hope, but none of them can satisfy my requirements.

I abandon Google and move to RunPod, a specialist GPU marketplace, and pay $10 up front. Same story, but playing out a little differently. Here availability is shown live on screen: a card indicates L4 GPU capacity exists, I select it, I mouse to the deploy button, and in the gap between those two motions somebody else takes it: “Instance not available.” One moment it is there. Then it’s not.

Now there’s no L4 capacity available so I try alternatives. A slot opens – a modest RTX 4000 Ada at $0.28 an hour. The pod queues, and the meter starts burning through credit while it initialises. Then the Whisper model downloads, the GPU reports in, everything works. I stop it to save money. When I try to start it again, the machine refuses: “There are not enough free GPUs on the host machine to start this pod.” Twice. The GPU I stopped has gone to someone else.

I start scrolling the price list, and here the story stops being a personal annoyance and becomes the thing we hear infrastructure leaders talking about. The catalogue runs from $0.17 an hour at the bottom to $7.89 an hour for a B300. The cheaper end – the GPUs a transcription server actually needs – is empty. The eye-watering price tags are available. If I were a big company – committed to renting GPUs at a discount or at a scale that meant I would make money from whatever I was doing with them – I would have leverage and I would have cards. As an individual who wants a few cents’ worth of compute, I have a spinning cursor.

The irony is that the hard part – the part I expected to be difficult – gave me no trouble. The server configured itself on the first attempt, the model loaded, the scripts worked. The software wasn’t the problem. What broke was simply renting the GPU.

I end the afternoon having transcribed nothing. My $10 now sits at $9.76. My Cline assistant that walked me through the latter parts of the process cost $0.24. It’s peanuts but it would have been nice to get a working transcription server out of it. I get to Sunday teatime and give up, get Cline to give me a transfer prompt to pick up the job later and chill out for the evening. 

But at least I got to see the GPU shortage from the bottom of the market, and a list of zones with nothing left to sell.



Source link