Ok stow away the hippie crap for a moment. The difference in scale between someone using their personal hardware to run models and the shit the hyperscalers are doing is not even fathomable.
My GPU uses a max of 250W if under full load, which it uses for a few seconds while running a simple request. A single Blackwell blows it out of the water, but those are not provisioned as single boards. They are provided in large racks, containing more CPUs, GPUs, RAM and Storage than i have ever possessed in my life, and i am at it since the early 90s. The difference in resource usage between those 2 scenarios isn’t even measurable because it is simply bonkers.
Ok stow away the hippie crap for a moment. The difference in scale between someone using their personal hardware to run models and the shit the hyperscalers are doing is not even fathomable.
My GPU uses a max of 250W if under full load, which it uses for a few seconds while running a simple request. A single Blackwell blows it out of the water, but those are not provisioned as single boards. They are provided in large racks, containing more CPUs, GPUs, RAM and Storage than i have ever possessed in my life, and i am at it since the early 90s. The difference in resource usage between those 2 scenarios isn’t even measurable because it is simply bonkers.