📊 Full opportunity report: Minerva. The opposite path. on ThorstenMeyerAI.com — validation score, market gap, and execution plan.
TL;DR
Italy’s Minerva project trained a large-scale Italian LLM from scratch, achieving impressive technical results but performing poorly on academic benchmarks. This reveals that scale alone may not ensure language-specific knowledge depth, challenging assumptions in European sovereign-LLM strategies.
Italy’s Minerva-3B, a large-scale, openly released Italian language model trained from scratch on 2.5 trillion tokens, scored only 4.9% on the INVALSI Italian school-exam benchmark, highlighting a significant challenge for European sovereign-LLM development.
Minerva was developed by Sapienza University of Rome’s NLP group, led by Roberto Navigli, using Italy’s CINECA supercomputing infrastructure and funded through national AI initiatives. It trained on approximately 50% Italian data, resulting in a 7B parameter model that outperforms comparable multilingual models on Italian benchmarks.
Despite technical success, Minerva’s performance on the INVALSI exam was near chance, a stark contrast to its impressive technical metrics. Researchers noted that while dataset composition is important, the overall size of data and parameters are more critical for complex language tasks, suggesting current investments may be insufficient for true language expertise.
Minerva.
The opposite
path.
Italy spent years building a European sovereign LLM from scratch. Then Minerva-3B scored 4.9% on the INVALSI Italian school exam.
Where AMÁLIA layered Portuguese specialization onto a multilingual foundation, Minerva trained from scratch on 2.5 trillion tokens with approximately 50% Italian content. Where AMÁLIA’s weights are not yet public, Minerva published weights, training data, and code as truly-open from day one. By every institutional measure, the Italian approach worked. But the empirical results contain a finding the press coverage has been quiet about — and it has implications that extend well beyond Italy.
Same problem. Opposite path.
European sovereign-LLM development has two primary architectural approaches. Italy chose from scratch with substantial native-language foundation. Portugal chose continuation pre-training of a multilingual model. The structural comparison surfaces what each commitment actually requires operationally.
The comparison is not “Italy did it better than Portugal.” Both projects respond to the same structural problem with different architectural strategies under different institutional and economic constraints. Italy’s national-AI investment is structurally larger by an order of magnitude — and Minerva is the visible artifact of that scale.

NVIDIA Certified Associate: Generative AI LLMs (NCA-GENL) (NVIDIA Certification Guides)
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
4.9% on INVALSI. The bitter lesson surfaces.
In June 2024, researchers evaluated Minerva-3B on the Italian school-exam benchmark. The result was unambiguous. This is not a critique of Minerva — it is a critique of the public discourse around what Minerva’s empirical results actually demonstrate.
350M to 7B. Four parameter scales, one architecture.
The Minerva model family covers four parameter tiers, each with specific training corpora. Each scale level reveals what the from-scratch path actually requires at different operating points.
Italian + English
100B English
~50% English
+ 200B code
Three answers. Same question.
Minerva, AMÁLIA, and OpenEuroLLM represent the three operational answers to the European sovereign-LLM question. Each makes different architectural and institutional bets. The strategic discourse benefits from treating all three as data points in the same empirical experiment.
Three standards the movement should adopt.
The structural critique generalizes beyond Minerva. The European sovereign-LLM movement benefits from internalizing these lessons across every subsequent national project. Italy modeled the openness standard; the movement should adopt it as norm.
Minerva is one valid answer to the European sovereign-LLM question. AMÁLIA is another. OpenEuroLLM is potentially a third. The strategic discourse benefits from treating all three as data points in the same empirical experiment rather than as competing national-prestige projects. More analysis like this is needed. Not less.
Implications for European Sovereign-LLM Strategies
The results from Minerva demonstrate that large-scale training from scratch, even with substantial native-language data, does not automatically produce deep language understanding. This challenges assumptions that scale alone guarantees language proficiency and suggests European projects may need to significantly increase investment to achieve country-specific knowledge depth, impacting future AI policy and funding decisions.European Sovereign-LLM Approaches and Challenges
Italy’s Minerva project represents a deliberate choice to train from scratch, contrasting with models like Portugal’s AMÁLIA, which extended multilingual foundations with smaller language-specific datasets. While Minerva achieved technical benchmarks and demonstrated the potential of European infrastructure, its poor performance on academic benchmarks reveals limitations in current scaling strategies. The debate around ‘from scratch or continuation’ is central to European AI policy, with Minerva exposing the need for potentially larger investments to reach desired language expertise levels.“The empirical results indicate that the European sovereign-LLM movement may need to confront a harder scaling reality than previously assumed.”
— Thorsten Meyer
Unanswered Questions About Scaling and Effectiveness
It remains unclear how different scaling strategies, dataset compositions, or continued training will impact Minerva’s future performance. The ongoing research aims to refine methodologies, but the precise investment threshold needed for true language expertise in European languages is still undetermined.
Future Research and Policy Directions for European LLMs
The Minerva team plans to continue iterating on training methodologies, including ongoing experiments with continual training and larger datasets. Policymakers and researchers will likely reassess investment levels and strategies, possibly prioritizing larger native-language data collections or hybrid approaches to improve language-specific performance. Further benchmarking and evaluation are expected in the coming months to determine the path forward.
Key Questions
Why did Minerva perform poorly on the Italian school exams?
Despite large-scale training, Minerva’s limited performance suggests that dataset size and parameter count alone are insufficient for deep language understanding, especially for complex academic tasks.
How does Minerva compare to other European language models?
Compared to models like Portugal’s AMÁLIA, Minerva trained from scratch on more native data and larger parameters but still struggled with language-specific benchmarks, raising questions about the effectiveness of current scaling strategies.
What are the implications for European AI policy?
The findings suggest that European sovereign-LLM projects may need to increase investment significantly to achieve the desired depth of country-specific language knowledge, possibly requiring larger datasets and more parameters.
Is the Minerva project still ongoing?
Yes, the team continues to refine their models and methodologies, with ongoing experiments aimed at improving performance and understanding the scaling requirements for native-language AI.
What does this mean for future AI development in Europe?
The results highlight the importance of strategic investment in native-language data and model scale, indicating that achieving truly proficient language models may require a reevaluation of current resource commitments.
Source: ThorstenMeyerAI.com