Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

Lets say an easy response takes 32k tokens in total, and to be generous, let's say it does 1 tok/s. This is already ~9 hours, and 32k reasoning tokens isn't even that much and as mentioned, K3 probably does the longest/most reasoning/thinking out of the available open weights models today, much like GLM. Just lowering that performance to 0.5 tok/s, would lead to ~18 hours for a simple prompt to receive an answer.

And then that's just for single prompts, what about agent harnesses, where before every tool call the model could reason a bunch?

I agree with you that real world results would be interesting, but I wouldn't hold my breath nor expect it to realistically be able to be useful. Still, people should try it, for science if nothing else :)

 help



Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: