Hacker News
new
|
past
|
comments
|
ask
|
show
|
jobs
|
submit
login
Jayakumark
on July 23, 2025
|
parent
|
context
|
favorite
| on:
Qwen3-Coder: Agentic coding in the world
What will be the approx token/s prompt processing and generation speed with this setup on RTX 4090?
danielhanchen
on July 23, 2025
[–]
I also just made IQ1_M which needs 160GB! If you have 160-24 = 136 ish of RAM as well, then you should get 3 tokens to 5 ish per second.
If you don't have enough RAM, then < 1 token / s
Consider applying for YC's Fall 2026 batch!
Applications
are open till July 27.
Guidelines
|
FAQ
|
Lists
|
API
|
Security
|
Legal
|
Apply to YC
|
Contact
Search: