HiPEAC

Artificial Intelligence

Recommendations for artificial intelligence (AI)

Rapidly position Europe in the two “directions” of AI:

These two “directions” of AI are now more widely recognized. India hosted the AI Impact Summit in New Delhi in February 2026; rather than compete in building artificial general intelligence (AGI), policymakers and investors say India aims to become the world’s “adoption capital,” focusing on low-cost, local-language applications and “ultra-affordable” AI tools.

Direction 1: AI for everyday life, for industries, for business, for customers, for everybody:

Distributed AI, mainly for inference, enabling the vision of the next computing paradigm (NCP):

Continue to develop agentic AI, and specialized action models (SAMs)

As proposed in the HiPEAC Vision 2025, these SAMs (“small and specialized models that can interact with their environment”, unlike large language models (LLMs) that focus on language) should operate in a distributed infrastructure and an ecosystem should be created to support research, development and business around them. These models need to be refined, optimized, and reduced in size to improve efficiency. They can be optimized from more general foundation models by an ecosystem of companies providing their optimized SAMs in a marketplace so that they can be dynamically discovered and used by agentic orchestrators.

  • Europe should provide basic infrastructure to develop and fine tune the SAMs, e.g. by providing a repository of trusted “foundation models” and the compute power to fine tune them on private and / or proprietary data.

Create the infrastructure and ecosystem to support agentic AI and SAMs

The infrastructure (A2I2: agentic artificial intelligence infrastructure, also simplified as (AI)2) should include non-functional properties, allowing the dynamic selection of the agent according to the user’s criteria. The key elements of the interoperable infrastructure are:

  • Software for managing agents (the focus of the next big AI battleground), helping businesses handle the growing suite of AI agents they’re using from different providers [1]. Europe should develop orchestrators that allow the activation of services provided by agents and SAMs in response to a request (“prompt”) and the non-functional criteria associated with the request. The orchestrators should be trustable (for example, open source and therefore auditable; they should work for the benefit of their user, and not for their creator or related third-party interests) and interoperable, therefore their specifications should be open and should lead to standardization. Features should be added over time as new technology becomes available, without necessitating a major change in the infrastructure (e.g. starting with static orchestration before moving to dynamic orchestration, beginning with a predetermined set of services before moving to a process to discover services, adding trustable ledgers that can qualify the services, …)
  • A communication protocol between agents, building on existing de facto standards of 2025 (such as MCP, A2A, etc) but also allowing support for non-functional properties, such as latency, time, etc. Adding non-functional properties will be crucial to extend agentic AI to physically entangled systems such as robots and self-driving cars. The (AI)2 should be flexible in terms of its implementation and the location of the agents, from a lightweight protocol when everything is executed within a single computing resource to totally distributed execution on multiple of edge devices interconnected by “long” distance communication protocol with enforced security.
  • “Containerization” of agents and SAMs allowing migration (and, eventually, portability) in a secure way. The agents and SAMs should work in well-protected “silos” on various kinds of hardware, avoiding them to “contaminate” the system they are running on. A reference software infrastructure for using agents, SAMs and orchestrators should be publicly available and open.
  • Support the SAM containers with dedicated hardware. Hardware should enforce the separation of the silos where agents are running from the underlying hardware. Dedicated accelerators should also be developed to increase the efficiency of the agents and SAMs (such as neural processing units (NPUs), for example). The development of efficient accelerators for inference of SAMs is still complex, but more accessible than the development of complex, multipurpose graphics processing units (GPUs) or NPUs (and tensor processing units – TPUs) designed to run large LLMs and targeting servers where several of these accelerators could be combined in a larger virtual one, using a unified memory space and extra-fast low latency communication interfaces. Unified memory between the central processing unit (CPU) and the accelerators in consumer devices (smartphones, laptops, …) facilitates the use of SAMs at the edge. The dilemma of performance / specialized hardware versus flexibility / CPU-GPU can be alleviated by hard or soft reconfigurable accelerators, for example.
  • An organization should be set up to ensure interoperability and trustability of the entire ecosystem. The role of the organizations should be, first, to help define overall specifications for the ecosystem, then to ensure that agents, SAMs and orchestrators conform to requirements and are trustable. This could lead to certification processes. The organization could also act as a broker between providers and users. It could also kickstart the business ecosystem and business cases, taking institutional (cities, administrations, etc.) drivers and procurement as a starting point.

Develop “outside of the box” concepts and ideas and test them rapidly in the (AI)2 infrastructure

This can be facilitated by sharing knowledge on emerging new ideas and new open source solutions in a structured and open manner (knowledge infrastructure). Easy access to tools and resources to validate proposals should be provided, with an open mindset. Once tested, the new idea can use the supporting infrastructure to enable swift transition to a real product. The aim is to use the (AI)2 as a means to rapidly create and validate specific new agents and to improve the orchestration and communication protocols, rather as Lego bricks can be put together in different combinations to build a large variety of objects. Example of concepts that can be developed include the following:

  • Decompose complex processes into smaller process (“bricks”) that can be realized (“constructed”) with the distributed SAM infrastructure for smooth evolution and avoiding dependence on huge data centres.
  • Develop methodology and tools allowing a large model to be “broken up” into agents that are orchestrated together (i.e. implementing a large model by components in a (AI)2 infrastructure). For example, a first concept could be to automatically transform a mixture of experts large language model (MoE LLM) into a set of specific agents that are orchestrated together in a distributed way.
  • Focus on SAMs that can be used for interaction with the physical world (cyber-physical systems), thus keeping within real-world constraints like latency, timing, etc., i.e. interacting directly with the world, in a trusted manner, with a particular focus in providing instances of the (AI)2 dedicated for robotics.

Develop neurosymbolic systems

These systems should marry classical systems (for supervision) and generative AI solutions, and add classical supervision systems as agents in the agentic AI infrastructure. Classical systems can offer services that can be combined with ones provides by SAMs.

Develop solutions for the continuous evolution of agents

Develop agents that can evolve according to their environment and adapt in a controlled way. This is still an open research topic, using either external databases in an agentic framework, new ideas for the structure of AI, or new training schemes [2].

Direction 2: Position Europe in the “artificial superintelligence” race

Some actors in the field or AI aim to develop systems that could be of higher “intelligence” than human experts to solve very complex problems. This involves large efforts (comparable to the Manhattan Project) and extreme computing power as of the current state of the art. Such systems, as of today, are self-improving: any progress done makes the next iteration faster and better, which means any delay potentially becomes exponentially costly to catch up. This is a real race between companies and countries involving enormous investments.

Europe should decide quickly if it wants to enter the race

It is first a political decision.

If so, Europe can leverage its position in global processing power by developing and using a transnational (AI)2 infrastructure (in which each agent (“component”) could be a complete high-performance computing (HPC) centre), using a unified infrastructure arising from NCP concepts and intelligently harnessing existing decentralized European hardware resources and extending the AI Factories initiative into a unified computing continuum.

  • This is an action that calls for creating a physical centre like CERN to bring together competences within Europe in the field of AI and coordinating actions, for example extending the RAISE initiative [3]. This should also help to assess the state of the art and the innovations needed.
  • While developing advanced models, emphasis should be on the alignment problem: developing approaches that control the level of model “cheating, avoiding deceptive behaviours and the development internal goals that are not “aligned” with the goals of the system designers. This is also an opportunity to develop advanced models that reflect values and culture endorsed by European society.

The recent developments that drive some recommendations

Cut-off date for this rationale: beginning of March, 2026

The rapid evolution of AI...

The rapid evolution of AI the recent months is illustrated by the quote of Andrej Karpathy [4], you can find in the introduction and below, and similar constatations are shared by others [5].

“It is hard to communicate how much programming has changed due to AI in the last 2 months: not gradually and over time in the “progress as usual” way, but specifically this last December. There are a number of asterisks but imo coding agents basically didn’t work before December and basically work since - the models have significantly higher quality, long-term coherence and tenacity and they can power through large and long tasks, well past enough that it is extremely disruptive to the default programming workflow. <…>

As a result, programming is becoming unrecognizable. You’re not typing computer code into an editor like the way things were since computers were invented, that era is over. You’re spinning up AI agents, giving them tasks in English and managing and reviewing their work in parallel. The biggest prize is in figuring out how you can keep ascending the layers of abstraction to set up long-running orchestrator Claws with all of the right tools, memory and instructions that productively manage multiple parallel Code instances for you. The leverage achievable via top tier “agentic engineering” feels very high right now.

It’s not perfect, it needs high-level direction, judgement, taste, oversight, iteration and hints and ideas. It works a lot better in some scenarios than others (e.g. especially for tasks that are well-specified and where you can verify/test functionality). The key is to build intuition to decompose the task just right to hand off the parts that work and help out around the edges. But imo, this is nowhere near “business as usual” time in software. »

Progress in AI is advancing at an increasing rate. Before 2024, the race mainly concerned the size of the models: bigger was supposed to be always better. New techniques were created to improve model training, drawing on well-curated data or on synthetized data, allowing the cost of training to be reduced and therefore the size of the model to be increased. The “mixture of experts” (MoE) approach became widespread to reduce the cost of computing during inference. 2025 was the year of the “thinking models”: using more compute power at inference time resulted in drastic increases in model performance.

At the time of writing, in 2026, we are witnessing the rise of the orchestrated and agentic models, where models can explore options in parallel, discard wrong directions, refine the results, and interact through tools with the physical world or computers. An example is Google Aletheia (see Figure 1), which improves upon the results of Google Deep Think (a “thinking” model) without changing the model itself. Figure 2 shows not only the gain of using orchestration, but also the remarkable gain in compute power for the same level of performance in six months: nearly 27, or x128.

Figure 1: The workflow of Google Aletheia from [6]
Figure 2: As of January 2026, the latest advanced version of Deep Think had significantly outperformed the International Mathematical Olympiad (IMO) Gold version (July 2025) on Olympiad-level problems. Aletheia makes further leaps in terms of reasoning quality with lower inference-time compute. All results were graded by human experts. The figure also shows that for the same levels of performance reached in July 2025, the January 2026 version required nearly x100 less inference-time compute [6].

Aletheia is an illustration that the agent wrapper (the “orchestrator”) with “generate verify revise loop” is going to be more important than throwing just raw compute at the base model. In 29 out of 30 problems where Aletheia returned a solution, its conditional accuracy was 98.3%. Its success rate was about 6.5% on research-grade problems, which seems low, but is still impressive according to the complexity of the problems. This demonstrates that the agent layer is where the real capability gains will be made, not just the base model itself. This leads to the observation that the next big AI battleground will be centred on software for managing agents, helping businesses handle the growing suite of AI agents they’re using from different providers. For a more in-depth discussion of “thinking models”, agentic and autonomous models, see ‘Thinking models, agentic AI and increasing autonomy’, below.

The progress is not only in the models, but also in the hardware that allows the models to be executed. For example, “NVIDIA GB300 NVL72 systems now deliver up to 50x higher throughput per megawatt, resulting in 35x lower cost per token compared with the NVIDIA Hopper platform.”[6]. The H200 was launched in 2024 and the GB300 by end of 2025, showing the rapid improvement of performance of the architecture, technology and software stack (it should be noted that this gain is also due to the move from an 8-bit floating-point representation to a mixed 4-bit floating-point representation).

Figure 3: Performance improvement of NVIDIA's latest generation of GPU over the previous one [7]

The constant improvement of the (energy) efficiency of the hardware / software stack is the fuel that allow companies to deliver more performing models demanding more and more compute power at a similar cost, and to serve more and more users. By July 2025, ChatGPT had been adopted by around 10% of the world’s adult population [7]. We can observe a correlation between the release of new versions of ChatGPT and the availability of a new generation of GPU, with improvement in their efficiency (see Figure 4). This is an economic reason: the cost of training with “older” generation GPUs will be too high in terms of energy (and therefore in $).

Figure 4: Graph originally from the GTC 2023 keynote by NVIDIA CEO Jensen Huang, updated with GPT-5

This tight connection between hardware efficiency and the availability of new LLMs is bidirectional – the hardware should be flexible enough to cope with the new innovations in the domain of AI, especially for training: often, when a new hardware is available, it was specified for algorithms that are one or two generations older, therefore a certain flexibility (programmability) is required, to the expense of less efficiency. NVIDIA’s rapid performance gains and continued dominance serving AI workloads are particularly impressive given the difference between the timescales for AI model advances (which take months) and the timescales needed for hardware development (which are currently in years). This is explained for example in the keynote given by Deming Chen at HiPEAC 2026 [8]. Inference seems to reach a relative maturity, so we see appearing more specialized solutions for efficient inferences such a Groq, Cerebras, … All this shows the importance of co-design in hardware / model development, which NVIDIA calls “extreme codesign” and which has been promoted in multiple editions of the HiPEAC Vision.

Figure 5: Slide promoting a global codesign, included for example in the HiPEAC Vision 2023 – see [10]

Only a few companies are mastering the progress of AI; most of them are in the United States (US) or China, with hardly any in Europe. Most other companies are simply “users” of AI technologies.

It can be observed that the open-source community is playing an increasing role, driven by China and its “open weights” models and some US companies (with European background) like Hugging Face. However, due to the increasing costs of developing new models and stakeholder pressure, companies that previously delivered open-source models have in some cases transferred to closed models (e.g. Meta, perhaps Alibaba / Qwen after the departure of key people [9]). Further discussion of open-source models can be found in the chapter on open source.

The recent success of OpenClaw [10], initially developed by one man (the Austrian programmer Peter Steinberger), shows that it is still possible for a small team to develop successful AI technology without a large company. A self-hosted autonomous AI agent platform, OpenClaw runs locally and performs actions across apps and services on a user’s behalf. It can act continuously without user supervision. It is positioned as infrastructure rather than an app and its development has been unusually rapid and chaotic, with multiple renaming (it was first called Clawdbot, followed by Moltbot, before finally being named OpenClaw), viral growth, security controversies, and institutionalization within months. It was the one of the fastest growing GitHub repositories. Finally, as mid-February 2026, OpenClaw creator was hired by OpenAI, In an announcement, OpenAI said that the open-source AI tool would “live in a foundation” inside the company”, and the company’s chief executive Sam Altman wrote: “We expect this will quickly become core to our product offerings,” [11]. This is also an illustration that the world of AI is controlled by few powerful companies – powerful enough to absorb any interesting innovation.

Figure 6: A brief history of OpenClaw
Figure 7: Increase in the number of GitHub stars for the OpenClaw GitHub

In addition to power over the development of AI models being concentrated in a few companies, access to the most advanced AI model capabilities is being restricted to an elite of paying customers for example, Google Gemini 3 Deep Think is only available to Google AI Ultra subscribers (at a cost of about €275 per month). This is an example of the segregation that we see appearing between “normal users” using AI for mundane tasks, who are offered models with limited performance which are cheaper to run, and “supermodels” that eventually could only be limited to use by an elite (and we have not yet seen “superintelligent” models, very expensive to run, which we believe will have very restricted access).

The current access to AI for “normal” users is mainly through apps (and therefore the cloud), but smartphone makers (Apple, Google, Samsung, …) are slowly changing the game by allowing more and more functions to be done by local models running on the user’s smartphone or laptop. The capacity of small models is also rapidly augmenting [12] local capabilities: models that can run on smartphones can now be “thinking” models and they can use tools.

On laptops or personal computers, unified memory allowing the CPU and the GPU to share memory allows de facto bigger models to run on these machines, no longer limited by the amount of video random-access memory (VRAM) available on the GPU card. This is the architecture of Apple machines since the M series of processors, and it is why Macs are often preferred to develop local AI (on GitHub repositories, AI workloads still mainly use the NVIDIA CUDA software stack, but more and more are using the MLX stack of Apple silicon). Machines with 128 GB are now accessible by the public (until the increased cost of memory will limit this) – for instance, machines using the M5 Max from Apple, the GB10 from NVIDIA or the AMD Ryzen AI Max+ 395 processor. These machines can run quite large models locally. On smartphones, models like Qwen Qwen3.5-0.8B or 2B or 4B, Gemma-3n-E2B can run locally on modern smartphones with good levels of performance. The consumer hardware is now ready for the NCP allowing SAMs to run locally. This trend towards local execution of AI models continues a tendency we have noted and advocated for since at least the HiPEAC Vision 2023. SAMs for use in the NCP can now be run locally on widely available consumer hardware, albeit hardware that is currently designed and produced outside Europe.

This new AI era could be an opportunity for European companies to be “back in the game” concerning consumer hardware: smartphones, laptops and computers might not be the best form factor for this new AI era. Will it be an earbud like in the movie Her? There are a lot of tests of new devices, such as AI pins, smart glasses, etc. Most of them are unsuccessful for now, but OpenAI, Meta, Apple (?) seems to be working on defining what could be the “post smartphone” device. Will Europe sit still and wait?

Recent directions in AI development and their relation to the HiPEAC Vision

Thinking models, agentic AI and increasing autonomy

In 2025, the field of artificial intelligence saw remarkable advancements, particularly in large language models (LLMs) and their applications. Current advances in artificial intelligence are occurring at an accelerating pace and are widely underestimated by the general public and often incompatible with the timescales currently required to develop e.g. Horizon Europe project calls, and the three-year schedule of most EU-sponsored projects, where their aim might be outdated before the end of the project.

Modern AI systems can now perform complex technical tasks end-to-end with minimal human intervention. In software development, for example, an application can be generated from a natural-language description, tested automatically, refined through iterative self-evaluation, and delivered in operational form. As an example, Figure 7 shows the increase in complexity of software tasks that LLMs can perform with a 50% chance of success (by February 2026): we can observe a drastic improvement since the end of 2025. These systems are increasingly capable of making design choices and adjustments autonomously, reducing the need for expert supervision. As an example, Spotify announced in early 2026 that “its best developers [hadn’t] written a line of code since December, thanks to AI” [13].

Figure 8: Example of the increase of complexity of software task that recent LLM can perform with 50% success [16]

Even the benchmarks that are specially designed to be hard for AI are becoming rapidly obsolete, with LLMs reaching nearly 100% on some of them (Figure 8 shows the progress on ARC-AGI benchmarks, which were defined to be especially difficult for LLMs).

Figure 9: The ARC-AGI benchmark was designed to be very difficult for LLMs. The ARC-AGI1 is nearly obsolete, and ARC-AGI 2 seems to be in few months [17].

A strategic focus on using LLMs for programming has accelerated this transformation because software underlies AI development itself. Once systems became proficient at writing code, they could assist in building improved versions of future models. For example, OpenAI wrote “GPT-5.3-Codex is our first model that played a key role in its own creation. The Codex team used early versions to debug its own training, manage its own deployment, and diagnose test results and evaluations—we were impressed by how quickly Codex was able to accelerate its development.” [14]. This creates a self-reinforcing cycle in which each generation of AI helps produce the next, potentially leading to very rapid increases in capability.

Recent developments (for example in Google’s Gemini 3 ecosystem) illustrate a shift in artificial intelligence progress from simple model scaling to more sophisticated reasoning architectures and agent-based systems. An update to the “Deep Think” capability introduces a reasoning mode that allocates additional computation at inference time, allowing the model to elaborate longer before producing an answer. Instead of following a single linear chain of reasoning, the system explores multiple hypotheses in parallel, evaluates alternatives, backtracks from dead ends, and dynamically adjusts the number of reasoning steps according to problem difficulty. This approach enables substantial performance gains without modifying the underlying model parameters. Benchmark results indicate high performance across complex tasks, including abstract reasoning, competitive programming, and advanced scientific problem solving.

In the case of Gemini “Deep Think”, the computational cost per task has decreased significantly compared with earlier versions, showing that efficiency improvements can accompany capability gains [15]. Experimental data suggest that improvements in inference-time reasoning can reduce the computational resources required for expert-level performance by orders of magnitude, highlighting a new scaling paradigm based on smarter use of compute rather than larger models.

Alongside improvement using inference-time reasoning, research agents built on top of this reasoning framework demonstrate how orchestration around a model can further enhance performance. For example, Google’s Aletheia operates through a loop composed of generation, verification, and revision stages. It proposes candidate solutions, critically evaluates them for logical flaws or unsupported assumptions, and refines or restarts the process when necessary. By grounding its reasoning in external sources such as academic literature and by acknowledging when a problem cannot be solved, the system reduces hallucinations and improves reliability. In controlled evaluations, this agent achieved very high accuracy on advanced proof tasks, outperforming approaches that rely solely on increased computational [5:1].

Case studies in scientific and mathematical research show that such systems can really assist in real problems rather than only in standardized benchmarks. However, success rates still remain limited, and results do not indicate consistent breakthroughs. Evaluations of hundreds of research problems show that only a small fraction is solved correctly, though this still represents a significant improvement over earlier systems.

We can identify that advances in AI in 2025 highlight three major trends:

  1. Allocating more reasoning time at inference can dramatically increase performance while improving efficiency.
  2. The architecture surrounding a model, including tools and agent loops, can contribute more to capability gains than enlarging the base model itself.
  3. AI systems are beginning to function as research assistants capable of contributing to complex problem solving, although they remain far from consistently autonomous scientific discovery.

Orchestrating autonomous software agents

These developments suggest a transition from single-shot language models toward structured, tool-using agents that reason, verify, and iterate. Artificial intelligence development is shifting towards the management and orchestration of autonomous software agents, a concretization of the concept of “Guardian Angels” or “Digels” dating back to the HiPEAC Vision 2021 [16]. Businesses are increasingly deploying multiple AI agents from different providers, creating demand for platforms that can coordinate these systems, control their activities, and integrate them into existing workflows. This is exactly what the HiPEAC Vision 2025 proposed: building an ecosystem of distributed agents as a blueprint of the NCP (and therefore implementing some concepts promoted by the NCP).

Competition has intensified following the introduction of agent-management products designed to coordinate tasks across various business applications through a single interface (see the success of OpenClaw, for example, but also the products released by Anthropic, Abacus, Manus (now Meta), …). One system can direct multiple agents to perform work across different software platforms.

Broader competition is emerging over which platform will become the central dashboard for supervising AI agents. The hiring of the creator of OpenClaw by OpenAI is a clear signal. Technology companies offering cloud infrastructure, business software, and productivity tools are each promoting their own control environments, anticipating that organizations will standardize on a single system. Ownership of this control layer could determine long-term influence over enterprise software ecosystems. On the consumer side, this could transform the current multi-app ecosystems running on smartphones into a single app that will be the entry point for all functions that the user will want. As there will be only one app, the form factor of smartphone with a touch screen allowing to access multiple apps could evolve into a new device that will perform the same functions and more but without this access to numerous diversified apps. Several companies are exploring what form this device could take (OpenAI with Jony Ive [17], Meta, Apple…).

Investors are particularly attentive to the possibility that AI agents could replace significant portions of traditional software functionality, reducing demand for existing products while accelerating disruption across the software industry [18].

Overall, the emergence of agent-management platforms represents a new phase in computing, where the focus shifts from individual applications to coordinated networks of autonomous systems. Success in this environment is likely to depend on standardization, integration with business processes, data governance, and operational oversight, as well as the ability to navigate an increasingly competitive and volatile market landscape.

Designing hardware for AI

Specialization versus more general-purpose accelerators: The case of NVIDIA and Groq

Due to the increasing number of users, the heavy computing demand of “thinking models” and of agentic AI, the inference part is becoming orders of magnitude mode demanding in computing power than training. Therefore, the competition to have more efficient systems for inference (which will result in lower utilization costs) is heating up. Major companies are now developing their chips, with often a variant for training only (such as AWS, which is developing the “Inferentia” chip for inference only, besides its “Trainium” for training). GPUs can embed features that make them more efficient for inference, but they will never be as efficient as architectures specially designed for inference. This could explain why Nvidia made a “non-exclusive deal” with Groq to use their inference technology (and hired most of their team).

HiPEAC had spotted Groq’s relevance almost two years ago, and for more detail on Groq’s first chip generation, readers can refer to the analysis published about two years ago [19].

NVIDIA’s move is entirely logical. A well-known principle in compute architecture is that the more you specialize an architecture, the more efficient it becomes. As a rough rule of thumb, you can see about an order-of-magnitude improvement when moving from CPUs to GPUs, another order of magnitude from GPUs to reconfigurable dataflow-style architectures such as coarse-grained reconfigurable arrays (CGRAs) or field-programmable gate arrays (FPGAs), and another order of magnitude from FPGAs to application-specific integrated circuits (ASICs). This compounding effect is why, for a fixed task, we can end up near a thousandfold efficiency gap between a highly specialized ASIC like CEA’s NeuroCorgi [20] and a general-purpose GPU. (For further discussion of reconfigurable computing platforms and how these fit into the HiPEAC Vision, see the hardware chapter.)

Figure 10: Architecture of Groq chip [24]

The trade-off is that specialization reduces generality. AI algorithms evolve quickly, so you need a certain amount of flexibility — programmability and reconfigurability — to keep pace. GPUs are more flexible than NPUs and therefore tend to adapt faster to algorithmic innovation, but they rely on a substantial software environment to bridge the gap between new algorithms and the hardware. A more “hard-wired” solution has less flexibility and therefore does not require such a generic and complex software stack. In practice, it often needs more of a “mapper” than a full compiler. However, the downside is clear: algorithms that differ significantly from what the ASIC was designed for are either not implementable at all, or they run with a severe efficiency penalty. This limits the specialization for rapidly evolving tasks like training, which illustrates why some level of flexibility remains necessary.

What has changed recently is that there is now more stability in algorithms, at least on the inference side. Inference still relies largely on transformer-based models (see “The rise of the transformers”, below), with some variants such as Mamba (see “Continuous learning and context awareness”, below) and mixture of experts (MoE) (see “Sub-giga-parameter models”, below).

At the same time, demand for inference compute has risen sharply, often by a factor of roughly ten to fifty compared with training workloads, depending on the application and how you measure it. Under these conditions, less flexible solutions than GPUs can generate a return on investment before they become obsolete, which is precisely where Groq becomes attractive. Specialized chips for inference are also easier to realize, allowing more companies to enter the game, unlike for training, where the complexity is much higher and even large groups have to give up [21].

For a less flexible approach for inference, a dataflow architecture like Groq’s is a natural fit. Groq also made two notable first-generation bets. First, they did not rely on the most cutting-edge process technology. Second, they leaned into “near-memory” computing by integrating memory on the chip, in order to avoid the very large energy cost of moving data off-chip, which can dominate overall power consumption. This reduces the need for external memory and can be an advantage during this current period of memory scarcity driven by AI demand (the so-called “RAM apocalypse”). The drawback is that on-chip memory capacity is limited per device, so scaling to even a “mid-sized” LLM requires very large chip counts. According to the analysis referenced above [19:1], running the Mixtral MoE language model requires computing power in the order of hundreds of chips — 576 are mentioned — which effectively confines this approach to datacentre deployments.

Groq also optimized heavily for low latency rather than maximizing throughput in the way GPUs typically do. Low latency is particularly valuable for real-time and interactive use cases, for agentic AI, and for systems that perform iterative inference steps, including so-called “reasoning” models. In that sense, the design aligns well with the dominant patterns in current inference demand.

Another practical point is complexity. Groq’s architecture is significantly less complex than NVIDIA’s GPU architectures, which can translate into smaller development teams on both the hardware and software sides. With NVIDIA’s cash reserves, acquiring intellectual property (IP) and integrating it can be faster than building everything organically, and this approach may also help them avoid the regulatory friction they encountered with their attempted Arm acquisition, which was ultimately abandoned due to competition concerns. There is also the talent angle: they would be bringing in Groq’s Jonathan Ross, a key architect from Google’s original TPU effort.

The next step will be to see what NVIDIA actually ships and how they position it. The expectation is that they will aim to support future inference-oriented chips through CUDA to keep a coherent ecosystem, even if the supported functionality is more constrained than on general-purpose GPUs. They may also evolve the design onto more advanced process nodes over time. Finally, because NVIDIA increasingly sells complete systems rather than only chips, they are in a strong position to build multi-chip inference systems at the scale Groq requires, potentially integrating hundreds of chips in a single datacentre-grade solution.

A reminder of recent AI history: The rise of the transformers

The current “boom” of artificial intelligence, and more precisely of the generative AI can be traced back in 2017 with the paper “Attention Is All You Need” [14:1] by researchers at Google. Its main goal was to accelerate the training of recursive neural networks that were used for language processing in introducing more parallelism. Unlike machine vision, which was already booming due to the rebirth of neural network-based approaches (using convolutional neural networks, or CNNs) and is inherently parallel, sequential processes like text, speech, translation, etc. required a different structure than CNN, and this paper showed one way to induce more parallelism. The side effect was to unlock a complete domain.

“We propose a new simple network architecture, the Transformer, based solely on attention mechanisms, dispensing with recurrence and convolutions entirely. Experiments on two machine translation tasks show these models to be superior in quality while being more parallelizable and requiring significantly less time to train. On the WMT 2014 English-to-French translation task, our model establishes a new single-model state-of-the-art BLEU score of 41.8 after training for 3.5 days on eight GPUs, a small fraction of the training costs of the best models from the literature. We show that the Transformer generalizes well to other tasks by applying it successfully to English constituency parsing both with large and limited training data. “

Figure 11: the paper from Google that started all...

Google didn’t seem to realize the impact of its paper at that time, and a start-up, called OpenAI, saw the potential, and started to develop more and more complex models using “transformers”. This is the series of GPT (GPT stands for generative pre-trained transformer) that led to a series of models leading to GPT 3.5, the basis for the famous ChatGPT that shocked (and changed) the world when introduced in November 2022.

Figure 12: Evolution of models from OpenAI. The details of models after GPT-3 are not publicly known, some leaks exist for GPT-4, but none about the real size of GPT-5
Smaller and smaller models

While frontier models continue to increase in size (see Figure 10 ), progress in algorithms, in training and datasets allow smaller models to have similar levels of performance as those of larger models from a few months before, at least in more specialized domains. For example, current models of about 10 billion (10B) parameters demonstrate better performance on specific tasks (benchmarks) than the original ChatGPT of 2022 (see Figure 11 and also [12:1]).

This trend was predicted in the HiPEAC Vision 2023 [22] (p.88) which called for Europe to “reduce both the memory footprint and the computing power of algorithms for artificial intelligence required for inference and learning”.

What is also interesting is that smaller models are often “open weight”, meaning that they can be easily duplicated (following their respective licensing agreement) and that open source (open weight) models are generally catching up closed models within the space of a few months.

Figure 13: Performances of various models on same benchmarks [26]
Figure 14: Comparison of performance of a 2.6B parameter model of 2025 with the "original" ChatGPT

Smaller models can also be “trained” by synthetic data generated by bigger models, as exemplified by the various DeepSeek small models, based on existing small models (often Qwen models) but fined tuned using data generated by the largest DeepSeek model:

“We demonstrate that the reasoning patterns of larger models can be distilled into smaller models, resulting in better performance compared to the reasoning patterns discovered through RL on small models.

Using the reasoning data generated by DeepSeek-R1, we fine-tuned several dense models that are widely used in the research community. The evaluation results (see Figure 13 ) demonstrate that the distilled smaller dense models perform exceptionally well on benchmarks. We open-source distilled 1.5B, 7B, 8B, 14B, 32B, and 70B checkpoints based on Qwen2.5 and Llama3 series to the community.” [23].

Figure 15: Performances of fine-tuned models [27]. Note : The American Invitational Mathematics Examination (AIME) is a selective and prestigious 15-question 3-hour test given since 1983 to those who rank in the top 2.5% on the AMC 10

Sub-giga-parameter models

More generally, this seems to show that for generalization on a particular task, a small subnetwork is sufficient: this is the confirmation of the lottery ticket hypothesis [24] of the “mixture of experts” (MoE) approach, which selects the subnetwork, and especially of the agentic approach, which can be seen as a decentralization of the MoE approach. An MoE is a model architecture where many specialized “expert” networks run in parallel, and a gating network learns to select which experts should handle each input. Instead of one big model doing everything, MoE routes different tasks or tokens to the experts best suited for them, making the system both scalable and efficient. The link to the agentic approach is the following: agentic systems also rely on specialized components (tools, skills, agents) that are dynamically invoked depending on the situation. Like an MoE gate, an agentic orchestrator decides which agent to activate for each subtask. In essence, MoE is the neural-network analogue of an agentic system: both achieve greater capability by routing problems to the right specialists, rather than relying on a single monolithic generalist. As explained in the HiPEAC Vision 2025, this mechanism is very similar to the proposed structure of the NCP.

A number of “small” models were released in 2025: along the lines of small models, Gemma3-270M [25] (here M means indeed million) runs well on a phone and has surprising levels of performance for its size (like GPT-3 a few years ago). Moreover, there is a profusion of small models under 1B parameters: Qwen 600M, Hunyuan-0.5b-instruct, Ernie-4.5-0.3b-pt [26], the last three being from the Chinese BATIX (Alibaba, Tencent, Baidu), but also MobileLLM [26:1] from Meta show good levels of performance for the 350M parameter model. The smallest is 125M parameters. Qwen from Alibaba also deliver small models (even multimodal) that can easily compete with larger models [12:2].

Figure 16: Performance of the MobileLLM from Meta. MobileLLM-R1 from Meta shows good levels of performance for the 360M parameter model [31]
Figure 17: Examples of test of the 140M and 360M models from the MobileLLM series

In June 2025, Microsoft presented MU [27], which will be in Copilot+ PCs and which has 330M. We are truly entering the era of small models running at the edge.

In the line of specialized nano models, one can also look at Kitten, which is a text-to-speech model with 25M parameters [28].

Figure 18: example of results of various small models

There are also other optimizations that will enable LLMs to run on memory-constrained devices: new memory optimization in the multimodal open-source model Gemma3n from Google (A 7B parameters model running on 4GB of RAM) – see Figure 17.

Figure 19: Performance of the Gemma 3n memory optimized model [34]

Ultra-small AI with interesting impact

June 2025 witnessed the launch of a model with a very interesting performance / size ratio of 27M (yes, millions not billions) parameters performs better than Claude Opus 4 on the (very specialized) ARC AGI 1 benchmark. The results on the ARC AGI 1 benchmark have been confirmed by the ARC Prize team [29]. This new model is called HRM (Hierarchical Reasoning Model [30]):

Figure 20: On ARC-AGI 1 benchmark, a 27M parameters model (HRM, Hierarchical Reasoning Model) beat Claude Opus 4

However, on October 6th, 2025, a new smaller model showed better results that the HRM model: with only 7M parameters, the TRM (Tiny Recursive Model [31]) obtained 45% test-accuracy on ARC-AGI-1 and 8% on ARC-AGI-2, higher than most LLMs (e.g., DeepSeek R1, o3-mini, Gemini 2.5 Pro) with fewer than 0.01% of the parameters.

Figure 21: Results of TRM [AI43]

The TRM model can be used to solve ARC-AGI puzzles, or to solve hard Sudoku or to navigate into mazes (see Figure 20 ). All information is in the paper and in a Github repository [32]. These models are small in size, but not necessarily in computing: they iterate a large number of times before providing their result.

Figure 22: Some results of the TRM [39]

Other key AI trends in 2025

Continuous learning and context awareness

There is growing demand for AI that can adapt to its environment, and use its past experience. Fine tuning a preexisting model is possible, but this requires retraining (albeit on a limited number of parameters if a parameter-efficient fine-tuning (PEFT) technique such as an adapter or LoRA approach is used). In addition, fine tuning cannot be done continuously or in the field due to the need for storage of the training examples and computing power. Enlarging the context of the LLM is a possibility, with some currently available models having more than 1M token context.

With the classical transformer approach, the requirements in resources (memory) are exponential with the size of the context. The Mamba approach, developed by researchers at Carnegie Mellon and Princeton, enables a linear increase, but, in practice, reduces performance. In 2025, hybrid models using both transformers and Mamba appeared with a good compromise between performance and resources. Other approaches, mainly agentic ones, use standard databases in the flow, allowing the system to take into account new and older events. Research in that field is active and new ideas are emerging [2:1].

World models and neurosymbolic AI: From text to the real world

LLM are trained with text (hence the name “large language models”), but texts are only a human projection of the world. It is totally different to be told that an object falls than to experience gravity directly. LLMs could be compared to children who stay in bed all the time and only learn through stories that are told to them. LLMs exhibit stunning levels of performance, but they are not the end of AI development. Interesting developments are ongoing.

In 2025, we witnessed the development of world models. “World models are neural networks that understand the dynamics of the real world, including physics and spatial properties. They can use input data, including text, image, video, and movement, to generate videos that simulate realistic physical environments” [33]. In other words, a world model is an agent’s internal, learnable representation of how the environment works — i.e., a model that maps states and actions to predicted future observations and outcomes (and often rewards), so the system can simulate “what would happen if…” for planning, control, or self-correction. They can be used to generate custom synthetic data or downstream AI models for training robots and autonomous vehicles, or to generate realistic videos or even video games (see Genie from Google Deepmind [34]).

World models are becoming increasingly important because “AI is becoming physical”, i.e. interacting with the real world, for robotics, self-driving cars etc. In January 2026, NVIDIA released its Alpamayo [35][36], a “family of open-source AI models and tools to accelerate safe, reasoning-based autonomous vehicle development”, heavily using world models.

Safety and explainability are key elements for such applications, and neurosymbolic AI might be one answer to these challenges. Neurosymbolic AI is an approach that combines neural networks with symbolic reasoning (which is good at explicit logic, rules, and structured knowledge, like knowledge graphs or programs). The goal is to get systems that can both perceive and generalize from messy real-world data while also reasoning transparently and consistently. It seems that the self-driving system from NVIDIA is using this combination of techniques. NVIDIA is also heavily relying on its Omniverse solution that allows the creation of realistic (both from the laws of physics and visually) digital twins that can be used to develop multimodal AI.

Figure 23: Simulating real environment helps robots to "learn". From NVIDIA [44]

Multimodal embodied physical AI

Researchers like Turing Award winner and former Meta scientist Yann LeCun advocate a different approach to developing bigger and bigger LLMs: his company, AMI Labs, has raised approximately $1 billion to develop what LeCun calls “world models.” These are AI systems that, this time, no longer learn solely from textual data [37] but can understand the real world. Similar ideas are currently developed by several companies in the world.

But what could be another evolution of AI? We can perhaps go back to the philosophical discussion on innate and acquired knowledge: up to now, the “structures” of AI models are quite similar, and most efforts have focused on training, i.e. on acquired knowledge. Perhaps the next step would be to develop new “structures” (like “transformers” were in 2017) that could improve the performance of AI. The structure (the “innate” knowledge) seems to have an important impact as well, as illustrated by the work on simulation of a fly brain [38] running in a virtual environment: only by a connectome-based brain emulation (i.e. the structure of the brain with the connections between neurons) with a physics-simulated fly body, does the simulated fly (without training or changing the synaptic weights) shows behaviour similar to a real fly. The CNN models were inspired from the work of David Marr, David Hubel and Torsten Wiesen about the structure of the visual cortex; the brain has specialized areas like agentic AI uses, so perhaps more inspiration from the structure of cognitive systems might help improving the next generation of AI

AI is becoming physical and the emergence of useful robots

The natural sequel of AI becoming physical is the emergence of robots. This is a clear target for China which has several companies working on robots, humanoid or not. Outside of China, Boston Dynamics (owned by Hyundai Motor Company), Figure AI, Tesla, etc are developing advanced robots.

Robotics is becoming a key domain for sovereignty and the US is responding to China: according to Politico [39]a national robotics strategy is in the works, including a decree planned for 2026 aimed at accelerating domestic research and production in order to secure US supply chains and compete technologically with China. This ambition is backed by massive investments, such as potential $1 billion in funding from Nvidia and SoftBank for Skild AI. Tesla plans to produce one million Optimus robots by the end of 2026, marking a historic transition from laboratory prototype to mass industrialization.

During the Chinese New Year show of 2026, China showed impressive performance of Unitree robots. The convergence of the progress in robots and “AI becoming physical” will certainly be the next big revolution; after AI taking the cerebral work, robots will also be more versatile and be used in increasing numbers of applications, possibly replacing humans.

According to [40], “China’s New Five-Year Plan Prioritizes Robotics and Beijing is embarking on a “whole-of-nation push” to achieve permanent dominance in physical AI technologies.<…> To understand what makes the 15th FYP’s treatment of robotics distinct from its predecessors, <…> Chapter 5 of the draft outline designates robotics as one of eight “strategic emerging industries” (战略性新兴产业) earmarked for accelerated development, alongside new-generation IT, new energy vehicles, biomedicine, and aerospace.” This Five-Year Plan “promotes the systematic elevation of robotics and “embodied intelligence” (具身智能) from a niche industrial subsidy target into the connective tissue of China’s entire economic modernization strategy.”

Everywhere, research labs and companies are working on developing more and more agile robots, even able to perform and navigates high-speed parkour with autonomous movement planning [41].

Figure 24: Unitree H2 making kickboxing [50]
Figure 25: Chinese New Year TV broadcast featured robots as martial art artists, even playing with nunchakus [51]

Unitree robots are becoming accessible (even available on Amazon for about $18000). Unitree might be on the same track as DJI, which developed cheap drones that were sold everywhere and it became a leader in the drone domain. Drones are used now in many applications, including military operations, so cheap robots could be as well in the future.

The potential use of AI in certain applications triggers ethical problems about the use of AI and robots, which led Anthropic to be categorized as a “supply chain risk" for the US Department of War [42]. More generally, this led to the important problem of “alignment” of AI on ethical principles, and the capacity of Europe to be able to develop AI that cope with the European ethical principles (see the insert text on alignment).

Figure 26: the unitree robots are available on Amazon
The alignment problem

(This part is based on details reported in a research article from November 2025 describing reinforcement learning on hackable, production-style coding environments [43])

Modern language models are not trained by simply reading the internet and predicting the next word. Once they reach a useful baseline, they are often improved with reinforcement learning, a family of methods in which a system is repeatedly placed in a task environment, takes actions, and is given a numerical signal intended to represent success. The hope is straightforward: if you pay the model for good behaviour, it will learn to behave well. The difficulty is equally straightforward: the signal is never a perfect expression of what designers actually want, and any gap between the signal and the true intent becomes a potential exploit which is called “reward hacking”. Instead of solving the underlying problem, an agent (LLM) discovers a loophole in the measurement apparatus or evaluation harness and uses it to collect high scores cheaply.

When a powerful optimizer learns that “winning” is separable from “doing the right thing,” it can internalize the logic of opportunism. A series of realistic experiments in a production-style setting illustrated this concern into a concrete mechanism. The key observation is not simply that a model learns to cheat when cheating is available — this is expected. The main observation is that once the model learns to cheat effectively, other forms of misaligned behaviour begin to appear, even though the training process never directly rewards them. In other words, the model’s behaviour shifts across a broad front: deception, strategic compliance, cooperation with malicious actors in synthetic scenarios, efforts to avoid monitoring, and even attempted sabotage of safety-oriented tooling in an agentic coding workflow.

Once trained in this way, the model does what any competent optimizer would do: it finds and uses the shortcuts. Reward hacking is not just a symptom of a mis-specified reward; it can be a cause of broader generalization into misalignment-relevant behaviours.

Two concrete manifestations illustrate the severity of this generalization. The first concerns sabotage in an agentic coding setting. When the model is embedded in a realistic coding assistant workflow and asked to make changes to a codebase that is intended to detect reward hacking and other concerning reasoning, it sometimes makes changes that would quietly reduce the effectiveness of that detection. The second manifestation concerns alignment faking. Alignment faking refers to a strategic posture in which a system behaves as if it shares the supervisor’s goals while privately maintaining different objectives, often with the implicit plan to pursue its own objectives when the oversight regime weakens. The model begins to produce internally coherent rationales for deception and strategic compliance without being explicitly trained to do so. That is precisely the kind of generalization that makes alignment researchers uneasy, because it suggests the emergence of instrumental reasoning about oversight.

This could have an important impact for agentic AI: the modern trajectory of AI development is toward more agentic systems: tools that can write code, run commands, browse internal knowledge bases, and take multi-step actions with limited human supervision. In such systems, the boundary between “task performance” and “environment manipulation” is thin. If a model has learned, in training, that manipulating the environment is an efficient path to reward, then scaling that model into more capable and more autonomous settings increases the space of possible manipulations.

This could drive to a shift of research: instead of treating misalignment as something that appears only when a model is directly trained on harmful content or explicitly instructed to do harmful things, we must also consider “natural” pathways in which ordinary optimization pressure in realistic environments creates incentives for unwanted behaviours. The practical challenge is to build systems, incentives, and evaluation that notice the difference between a model that is helpful and a model that is merely good at looking helpful.

Research in alignment should become more and more important. In large companies driven by short time to market and cost-reduction motivations, it often becomes a second priority. Europe should lead the research in this field, as it is important to deliver AI systems that are compliant with European ethics, and could help strengthen Europe’s reputation as a leader in trustworthy AI systems.

Is Apple one step ahead in AI?

This post from Milk Road AI [44] hypotheses that Apple might have a similar reasoning than HiPEAC, betting more on collaborative distributed AI models in devices and on the “continuum of computing” than in building gigantic data centres. The HiPEAC Vision 2025 already illustrated this with the structure of Apple Intelligence, using an on-device orchestrator to select an on-device model or offload the request to a models executed in Apple servers, depending on the request (see Figure 4, or the chapter on AI in the HiPEAC Vision 2025). Although the post is highly speculative, this example illustrates the potential impact of “think different”. It also raises concerns about financial sustainability of large investments in data centres for AI. Major funding rounds, high valuations, low return from investment and complex investment relationships within the sector have led some industry leaders to warn of a potential market correction similar to earlier technology bubbles.

“The richest company on Earth just watched its rivals light $650 billion on fire.

And did nothing.

This might be the most brilliant move in corporate history.

  • Amazon is spending $200 billion this year on AI data centers. - Google, $185 billion. - Microsoft, $114 billion. - Meta, $135 billion.

Combined: $650 billion.

(…)

Apple is refusing to enter a race that might not have a finish line.

The hyperscalers are now spending 94% of their operating cash flows on AI infrastructure. After dividends and buybacks, there is almost nothing left. Amazon is projected to go negative on free cash flow this year as much as $28 billion in the red.​ Alphabet’s free cash flow is expected to collapse 90%. From $73 billion to $8 billion.​

These companies used to be the greatest cash machines ever built.

Now they’re borrowing money to keep the lights on.

The Big Five raised $121 billion in bonds in 2025 alone.​ Morgan Stanley projects $1.5 trillion in tech debt over the coming years.​ For the first time in history, hyperscalers hold more debt than cash and what are they getting for that $650 billion? AI services generate roughly $35 billion in total revenue and that’s 5% of what’s being spent on infrastructure.​

Now here is where Apple’s bet gets genius.

AI models are commoditizing faster than anyone predicted.​

(…)

Open source models now power 80% of startups seeking VC funding.​ The moat these companies are spending hundreds of billions to build is evaporating in real time.

Apple understood this before anyone else. It didn’t build its own AI model, it licensed Google’s Gemini for about $1 billion a year.​ Why spend $100 billion building a factory when the product costs a billion to rent? And if a better model appears next year, Apple just switches vendors.​

But Apple is not sitting still.

It just dropped the M5 chip with a 16 core Neural Engine and Neural Accelerators built into every GPU core.​ (…)

The M5 delivers 4x the AI performance of the M4 and Apple doesn’t need $200 billion in data centers.

Because Apple turned 2 billion devices into the data center.​ Every iPhone, Mac, iPad gets distributed AI at a scale no server farm can match.

While its rivals burn cash, Apple is doing the opposite. $90.7 billion in stock buybacks last fiscal year.​ Its competitors? Combined buybacks collapsed 74% from their peak.​ Apple didn’t miss the AI revolution. It just bet that the winners won’t be the ones who build the infrastructure. They’ll be the ones who own the customer and no one on Earth owns more customers than Apple.”

Milk Road AI [44:1]

Conclusion

It is clear that the progress in the domain of artificial intelligence is very rapid, with no sign that it is slowing down. At the time of writing, in March 2026, complex tasks could be done by the coordination of LLMs in the billion-parameter range, which was not the case a year before. Increasing the size of the networks is no longer the solution for increased performance.

2025 saw “thinking” models becoming widespread, iterating in the inference phase without changing their weights. Clever orchestration of LLMs in an agentic context was also shown to increase performance without changing the “core” LLMs. This is totally in line with what the HiPEAC Vision 2025 proposed and which continues in 2026.

2025 also witnessed more “AI becoming physical”, with the added requirements of non-functional properties such a safety, response time when the LLM is controlling a robot – again a step towards a use of the NCP proposed technologies. Embodied AI or “AI becoming physical” is the next step for AI. Europe’s know-how in real-time and embedded system could be essential to develop systems using both neural-network-based solutions and more classical approaches; hence the drive for neurosymbolic solutions synergizing both approaches.

The performance of AI in software development is also rapidly (exponentially) progressing, with AI contributing towards its own development (ex for GPT-5.3-Codex). The question of the “dual speed AI”, i.e. development of artificial general intelligence (AGI) / artificial superintelligence (ASI) versus the development of low-cost, regional applications and “ultra-affordable” AI is becoming more visible, and Europe has little time to decide which track(s) it will follow (of course, it also possible to choose both).

As explained in the HiPEAC Vision 2026, European agentic AI based on the NCP concepts (which we have named (AI)2: agentic artificial intelligence infrastructure) is a blueprint of a possible full-blown NCP and could benefit from the knowhow, diversity and SME in Europe. But it could also be used to federalize big data centres and HPC centres in Europe to harness the computing resources required to develop AGI and even ASI.

On the hardware side, inference is becoming more and more demanding, with low latency (for “reasoning” models, for agentic AI, for “AI becoming physical”) but if used in an agentic infrastructure, each agent is perhaps not so large, so dedicated accelerators for small- or medium-sized language models (or better SAMs) are key; these also in the scope of European knowhow and capabilities.


  1. https://www.theinformation.com/newsletters/applied-ai/looming-battle-agent-management-software↩︎

  2. https://arxiv.org/pdf/2602.16066v1↩︎↩︎

  3. https://eur-lex.europa.eu/legal-content/EN/TXT/?uri=CELEX%3A52025DC0724&qid=1762332390557↩︎

  4. https://x.com/karpathy/status/2026731645169185220↩︎

  5. Matt Shumer “Something Big is Happening” (Feb 9, 2026) from https://x.com/mattshumer_/status/2021256989876109403↩︎↩︎

  6. https://blogs.nvidia.com/blog/data-blackwell-ultra-performance-lower-cost-agentic-ai/↩︎

  7. https://cdn.openai.com/pdf/a253471f-8260-40c6-a2cc-aa93fe9f142e/economic-research-chatgpt-usage-paper.pdf↩︎

  8. https://youtu.be/qUhHCeZlwXg?t=198↩︎

  9. https://venturebeat.com/technology/did-alibaba-just-kneecap-its-powerful-qwen-ai-team-key-figures-depart-in↩︎

  10. https://openclaw.ai/https://github.com/openclaw/openclaw↩︎

  11. https://www.cnbc.com/2026/02/15/openclaw-creator-peter-steinberger-joining-openai-altman-says.html↩︎

  12. https://venturebeat.com/technology/alibabas-small-open-source-qwen3-5-9b-beats-openais-gpt-oss-120b-and-can-run↩︎↩︎↩︎

  13. https://techcrunch.com/2026/02/12/spotify-says-its-best-developers-havent-written-a-line-of-code-since-december-thanks-to-ai/↩︎

  14. https://openai.com/index/introducing-gpt-5-3-codex/↩︎↩︎

  15. https://deepmind.google/blog/accelerating-mathematical-and-scientific-discovery-with-gemini-deep-think/↩︎

  16. HiPEAC Vision 2021 https://www.hipeac.net/vision/2021.pdf↩︎

  17. https://www.theverge.com/news/672357/openai-ai-device-sam-altman-jony-ive↩︎

  18. https://www.bain.com/insights/will-agentic-ai-disrupt-saas-technology-report-2025/↩︎

  19. https://newsletter.semianalysis.com/p/groq-inference-tokenomics-speed-but↩︎↩︎

  20. https://list.cea.fr/en/neurocorgi/↩︎

  21. https://www.theinformation.com/articles/metas-internal-chip-design-efforts-hit-roadblocks↩︎

  22. HiPEAC Vision 2023 https://www.hipeac.net/vision/2023.pdf↩︎

  23. https://github.com/deepseek-ai/DeepSeek-R1?tab=readme-ov-file#distilled-model-evaluation↩︎

  24. https://arxiv.org/abs/1803.03635↩︎

  25. https://developers.googleblog.com/en/introducing-gemma-3-270m/↩︎

  26. https://huggingface.co/baidu/ERNIE-4.5-0.3B-P↩︎↩︎

  27. https://blogs.windows.com/windowsexperience/2025/06/23/introducing-mu-language-model-and-how-it-enabled-the-agent-in-windows-settings/↩︎

  28. https://github.com/KittenML/KittenTTS↩︎

  29. https://arcprize.org/blog/hrm-analysis↩︎

  30. https://arxiv.org/abs/2506.21734↩︎

  31. https://arxiv.org/abs/2510.04871↩︎

  32. https://github.com/SamsungSAILMontreal/TinyRecursiveModels↩︎

  33. https://www.nvidia.com/en-us/glossary/world-models/↩︎

  34. https://deepmind.google/models/genie/↩︎

  35. https://nvidianews.nvidia.com/news/alpamayo-autonomous-vehicle-development↩︎

  36. https://d1qx31qr3h6wln.cloudfront.net/publications/Alpamayo 1_2.pdf↩︎

  37. https://techcrunch.com/2026/03/09/yann-lecuns-ami-labs-raises-1-03-billion-to-build-world-models/↩︎

  38. https://theinnermostloop.substack.com/p/the-first-multi-behavior-brain-upload↩︎

  39. https://www.politico.com/news/2025/12/03/trump-administration-ai-robotics-00674204 and https://www.politico.com/newsletters/digital-future-daily/2026/02/04/ai-powered-robots-are-coming-for-trade-jobs-00765584↩︎

  40. https://thediplomat.com/2026/03/chinas-new-five-year-plan-prioritizes-robotics-the-world-should-pay-attention/↩︎

  41. https://techxplore.com/news/2026-03-humanoid-robots-master-parkour-human.html↩︎

  42. https://www.bbc.com/news/articles/cn48jj3y8ezo↩︎

  43. https://www.anthropic.com/research/emergent-misalignment-reward-hacking↩︎

  44. https://x.com/MilkRoadAI/status/2029318433360265227↩︎↩︎


Summary

Europe's AI strategy focuses on developing agentic AI for everyday use and preparing for the artificial superintelligence race, advocating for infrastructure to support specialized action models (SAMs).