Nuxt @nuxt.com · Feb 26

We now run evals to help you know what AI models perform best to write Nuxt code. nuxt.com/evals

70 likes 8 replies

?

Replies

Lukas Trumm · Feb 27

Awesome! My experience is that for opinionated rules, its still better to mostly include the rules and fine tune them for better results. Might be short, models understand. Here is my own take on evals (focusing on opinionated Vue rules and Claude Code). vue-nuxt-rules.lttr.cz/rule-evals

Nathan Chase · Feb 26

I wonder if this means that: a) Nuxt's error handling docs are insufficient b) The error handling itself is incomplete, or could be better c) Opus just has a hard time dealing with error handling, generally d) All of the above I recently implemented @hrcd.fr's evlog w/ Opus's help. So far so good.

Stephanie · Feb 26

That’s basically my experience prompting different models with different Nuxt issues. Any thoughts on adding OSS models as well? Maybe in a separate list from the commercial options though?

dame · Feb 26

This is actually awesome

Adityawarman Dewa Putra · Feb 27

Gemini 3.1?

Florian van der Galiën · Feb 26

What a brilliant idea! 🤩

Loic · Feb 26

Awesome! Any chance you could add open source Mistral Devstral2 to the list? mistral.ai/fr/news/devs...

Zoe 🏳️‍⚧️ · Feb 27

I’d be interested in running evals for a ton of OSS models. What does the token count look like for an entire run?