AI in the real world: from chatbots to actions

Just a couple of years ago, the main battleground in the artificial intelligence industry was text, images and code. Today, the focus is shifting towards the physical world, according to experts at NIFOROSERNO. The new generation of models is learning not just to answer questions, but to pick up objects, open drawers and move things around – in other words, to act in ways that previously only humans and specially programmed robots were capable of.

Multimodality as a new foundation

Multimodality – the ability of systems to simultaneously process text, images, sound and sensor data – has been the key driver of this shift, according to Niforoserno Canada. Whereas previously a model could only describe what it saw in an image, it is now capable of linking that perception to a specific physical action. The robot ‘sees’ a cup on the table, understands the command ‘move it further to the left’ and carries out the task, drawing on a single model that integrates vision, language and motor skills.

This approach fundamentally changes the architecture of the systems. Instead of a set of highly specialised programmes, developers are creating universal models that are trained on vast datasets of video footage showing how people manipulate objects in everyday life and in industrial settings.

From content generation to object manipulation

For a long time, generative AI was associated primarily with the creation of content – text, images and music. A new phase in the technology’s development is shifting the focus towards actions. Models trained on data relating to movement and object grasping are already capable of performing basic everyday tasks: opening a cupboard door, picking up a tool, or moving a box on a conveyor belt.

This does not negate previous achievements in content generation, but adds a new dimension to these systems – the ability to influence the physical environment directly, without human intervention, which was previously an essential link between the AI’s decision and its execution.

Robotics is gaining an intelligent core

Niforoserno emphasises that this marks a paradigm shift for the robotics industry. Previously, robots were programmed for specific tasks and struggled to cope with unexpected situations. Now, the focus is on a universal model capable of adapting to new conditions – such as a different arrangement of objects, changed lighting or an unfamiliar object.

This is precisely why major technology companies and start-ups are increasingly investing in the integration of large language models with physical platforms – manipulators, mobile platforms and humanoid robots.

What this means for business and everyday life

The practical application of such systems is already becoming apparent in logistics, warehouses and production lines, where robots with an intelligent core will be able to adapt flexibly to changing tasks without the need for reprogramming. In the future, similar technologies may also find their way into homes – in the form of domestic assistants capable of carrying out simple tasks.

According to experts at Niforoserno Digital Enterprise, despite impressive progress, the path from laboratory demonstrations to widespread use will not be a quick one. Questions remain regarding safety, reliability in unpredictable conditions and the cost of the equipment. Nevertheless, the direction has been set – artificial intelligence is increasingly moving beyond the screen and learning to operate in the physical world.

Related Posts