Ive been working with smaller LLMs for a while now on my old rig but its definitely showing its age with some of these newer quantization methods. I have about 500 bucks to spend and im trying to decide if I should hunt for a used 3090 or just grab a new 4060 Ti 16GB. My logic was that VRAM is king for inference but the memory bus on the 40 series is so narrow it feels like a trap. I need something decent by next month for a local fine-tuning project and im based in the US. Do you think the extra VRAM on the 4060 Ti makes up for the slower speed or should I just gamble on a used card...
Honestly, skip the NVIDIA GeForce RTX 4060 Ti 16GB and hunt for a used NVIDIA GeForce RTX 3090 24GB GDDR6X. The memory bus bottleneck is real; you'll be much happier.