Why aren't more LLMs quantized to INT8 W8A8 for RTX 3090?

RTX 3090 is the second most used GPU by LLM enthusiasts and has native INT8 tensor cores, yet most users default to FP8 or smaller quants. The Reddit discussion explores why W8A8 models aren't more common.
1 source
AI Models by email
Get an email when there's news on AI Models
No news that day, no email.
More stories today
- AI startup Wonderful raises funds at $5 billion valuation
- NYC bans AI use for students until high school
- Qwen Live Host v0.2.0 released
- Filevine launches AI citator and hallucination checker in LOIS
- AI billionaires fund ad blitz as data center opposition hits 61%