get_flat_ccs_offset() reads the base of the flat CCS storage from the
hardware, scales it by the number of enabled L3 nodes, and rounds the
result up to 128K. Everything below that offset is then ...
They are not magic. You are still using electricity, and for the same amount of work that cloud providers do you will consume the same amount. If every one of the billions of people using chatgpt or whatever buys rtx 5090 and makes it churn 24/7 for their “local” ai then nothing will change for the environment.
It will still run necessary calculations much slower than a gpu, so you will need to run it much longer for the same result, which is not only inconvenient but will likely use more power in the end anyway.
use Local LLMS ;)
They are not magic. You are still using electricity, and for the same amount of work that cloud providers do you will consume the same amount. If every one of the billions of people using chatgpt or whatever buys rtx 5090 and makes it churn 24/7 for their “local” ai then nothing will change for the environment.
no you won’t, datacenters inherently use far more electricity because they need enormous and intensive water cooling systems to cool the equipment
99% of consumer gpus, even the high end, are air cooled with tiny fans
But isn’t it as bad as running a graphically intensive AAA game on your computer
That’s still just moving the impact from data centers to end users’ machines.
Gamers (especially those that own 5090) are a tiny minority compared to llm users.
CPU?
This will just waste even more energy since cpus are less efficient at this kind of task.
LLMs require insane amounts of compute, period. The hardware to do that needs energy and emits heat, there is no way around.
What if it the CPU was energy efficient?
It will still run necessary calculations much slower than a gpu, so you will need to run it much longer for the same result, which is not only inconvenient but will likely use more power in the end anyway.