Hacker Newsnew | past | comments | ask | show | jobs | submit | enraged_camel's commentslogin

We use Haiku 4.5 inside our product. It continues to be absurdly capable for converting natural language to structured JSON based on a set of fairly complex business rules.

bro why. its literally the most overpriced model in existence right now. i could name about 10 models off the top of my head that would be better and cheaper

We tried Luna and it scored way lower in our evals. Muse also. We haven't had a chance to test others.

How much time were you able to put into tuning your prompts? And was it worse on all fronts (cost, latency, accuracy) or just some?

Ah, so you didn't read the article.

>> On our benchmarks, Claude Opus 5.5 leads in agentic coding, computer use, and knowledge work. That said, at these levels of capability we’ve found that benchmark margins have become a less reliable guide to real-world differences. In our own use, the gap between Opus 5.5 and Claude Fable 5.1 is narrower than these scores suggest.


You didn't read my question, bc that excerpt doesn't answer, nor do they demonstrate

> how would this alleged difference (most likely bs) actually show up in reality?

Furthermore: so they admit it's bs but still placate it like its the next biggest thing ever ... alright

All I'm saying is I refuse to buy into it anymore – yet many on here still do, including ... you?


At the end of the post they said Sonnet 5.5 and Haiku 5.5 are coming soon.

I'm confused. Why OpenAI and not Anthropic? I don't see anything here that is specific to OpenAI.

It's not just OpenAI. It can be any frontier-level lab that has more funding than Jev.

>> Why have we not seen an improvements in products?

I don't know what products you're referring to. At work we rewrote our platform from scratch using AI and our users have been raving about the massive improvements we made and the new features we added (that we previously hadn't been able to due to tech debt and being short-staffed).

We had a major account that dropped us last year because the main feature they were using didn't do what they needed and it had a lot of rough edges. We thought we had lost them forever. Well, yesterday our senior account exec convinced them to sit in on a demo of the newly built product and they were floored. They loved it so much that they're coming back.

Outside of work, for my side business, my users have been raving about the features I have been shipping for the past year. Previously I would work on it casually and mostly address support issues that came up, but shied away from rocking the boat too much with anything too ambitious because it's just me and my spare time is limited. But with the help of AI the product is now more performant, more user friendly, and has a lot of valuable features it was missing before. I've been ten times more ambitious. I know several solo founders like me who have similar stories.


Can you link one of them? Not that I don't believe you, I'm just curious

This thing is DOA. They compared 4.7 xhigh to 4.6 high to make it look like it improved. The reality is pretty bad: https://x.com/chetaslua/status/2102087511367618942

Astra fails in similar ways, and at similar frequency, as GPT 5.6 Sol does. It often goes way out of scope, or just stops prematurely, or tries to find odd and even dangerous workarounds when it gets stuck.

It's phenomenal at computer use and 3D stuff. I've been using it less and less for coding.


Same, Astra is extremely RL fried, and nobody is talking about it. I used Astra for a few days on my personal project, and load times went from less than 3 seconds to almost 30 seconds because it kept using the wrong sync primitives and bad architecture overall.

Huh, I've had a totally different experience. I've used it extensively, maxing out the 200€ plan on personal projects and it's the best model I've ever used, so easy and pleasant to use. It's great for frontend design and using it in Rust I've had Coming from Opus 5, it's a breath of fresh air.

Same experience. Astra is on par with or better than Fable 5.1 with a lot more usage on the plans. It has been an extraordinary experience using it so far. 5.6 Sol was very good and Astra is a large upgrade in quality.

Same. GPT-6 has been a huge breath of fresh air for me. Fixed 80% of the issues I was having with Sol. I just gave it the same task I gave to Sol a few months ago, and it knocked it out of the park comparatively.

LLM's introduces problems, and it finds them in its own internal thinking. But instead of actually modifying the previous generated answer to fix the real issue, it adds another layer to deterministically guard around it, greatly expanding the scope of the fix. This scales with effort, and the result is spaghetti and with a side of bugs.

Best to stick with a high end model + low effort, do a manual pass on high effort and fix the bugs you know are reachable.


It's really interesting how different the experience people have is with these models. I tried Codex with whatever they had before Sol and then with Sol, and just kept going back to Claude Opus/Fable because they were better at the coding work I was doing. Despite getting annoyed at the way it replied/wrote, it was just much better. Astra is the first one that feels as good as Fable to me, and it's much less annoying in its replies. I still don't think they have anything I'd want to drop down to like I can drop down to Opus though.

yeah I see this in these threads, I'm guessing the user prompts are the actual wildcard, it has been for my use thats for sure. edit: I wonder if gemini is somehow training me to like it more lol


Bus schedules are not a good comparison because they have to fit a ton of information into a single page. And I disagree that they are unreadable. Speaking as someone who commuted by bus for years in many different cities, they tend to be quite easy to read because the tabular layouts make them very easy to scan quickly. Once your eyes anchor on the page, you can find what you're looking for without effort.

Yeah. Let's not forget that just a year ago, those of us in tech could not conceive of developers getting replaced by AI. Things have come so far since then however that there are multiple studies showing junior developer hiring has slowed down to a crawl.

Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: