Some_Emo_Chick@lemmy.world to Technology@lemmy.worldEnglish · 6 days agoGenerative AI Is an Engineering Disaster. A shockingly inefficient trillion-dollar project.www.theatlantic.comexternal-linkmessage-square208linkfedilinkarrow-up1797arrow-down119cross-posted to: usa@midwest.social
arrow-up1778arrow-down1external-linkGenerative AI Is an Engineering Disaster. A shockingly inefficient trillion-dollar project.www.theatlantic.comSome_Emo_Chick@lemmy.world to Technology@lemmy.worldEnglish · 6 days agomessage-square208linkfedilinkcross-posted to: usa@midwest.social
minus-squareBrett@programming.devlinkfedilinkEnglisharrow-up4·6 days agoIs that quantized? 4 bit Qwen 3.6 can get 22tps on a 1060.
minus-squareAsafum@lemmy.worldlinkfedilinkEnglisharrow-up3·6 days agoIt’s the q4 quantization, but it requires 20+GB vram and my 5080 only has 16
minus-squareDamage@feddit.itlinkfedilinkEnglisharrow-up3·5 days agoMy framework 13 with shared RAM runs qwen quite well
Is that quantized? 4 bit Qwen 3.6 can get 22tps on a 1060.
It’s the q4 quantization, but it requires 20+GB vram and my 5080 only has 16
My framework 13 with shared RAM runs qwen quite well