News
2026-09-24
Human rights: part of professional training?
By: Alejandro García Suárez – Advisory Office of Communications
▲ Back to Top
Just 24 hours later the presentation of ChatGPT-4o, Google has matched and raised the stakes on Tuesday with the most advanced version of OpenAI's conversational robot, presenting similar improvements for its search engine, which are already starting in the United States and will be rolled out to the rest of the world.
The new search platform replicates the capabilities of what the company calls "agents," which can plan and execute actions on behalf of the user, but it humanizes the process to emulate an interaction with a person. Gemini, as the multinational's artificial intelligence and search engine are called, can be interrupted to redirect the conversation, and the mobile phone's camera becomes its eyes to describe what it sees, solve problems it observes, or pinpoint the location of an object it has registered during the conversation. Where did I put my keys? What's the solution to this problem? What is this? Ask Gemini.
Google has pulled out all the stops to counter OpenAI and fight for its dominance in the search arena. The company's CEO, Sundar Pichai has taken over the presentation The latest advances in artificial intelligence will be discussed this Tuesday at the annual edition of Google I/O in Mountain View, California. It will be applied to all products (Gmail, Photos, Drive, Meet, and any Workspace tool), but especially, according to Pichai, to the platform that is its stronghold: “The most exciting transformation with Gemini, of course, is in Google Search. We radically changed how it works.”.
“Gemini can maintain a personalized and interactive conversation, mixing and matching inputs and outputs,” Pichai explains regarding the humanization of the search engine's interaction, which moves away from a linear process (successive queries and responses) to emulate a relationship similar to a personal one. These are skills they already demonstrated to agents in Las Vegas last April, during the Google Next, where the robots that plan and execute actions on behalf of the user were launched.
“These are intelligent systems that demonstrate reasoning, planning, and memory. They are capable of thinking several steps ahead and working across all programs and systems, or of doing something on behalf of the user and, more importantly, with their supervision. We are giving a lot of thought to how to do this in a way that is private, secure, and works for everyone,” the executive explained in response to the ethical risks identified by the company’s own research group (DeepMind).
The conventional search engine, which returns web pages more or less related to the user's query, becomes a thing of the past with Gemini. Liz Reid, director of Google Search, says that although this tool has been "incredibly powerful," it requires "a lot of work," referring to the task of refining the descriptors and filtering the relevant information from the thousands of results obtained. "Searching has been through one question after another," she admits.
The new skills, he explains, understand “what you really have in mind,” contextualize, know the interaction point, and “reason” to offer a result that combines findings from various domains and presents a plan and advice. He illustrates this with a practical example: while the traditional search engine was asked for restaurants in the area, thanks to the AI Overview From Gemini, you can now request “a place to celebrate an anniversary” and the search engine offers different categories of plans, prices, locations, and suggestions. Or you can also provide a complex travel program for a family of several members with different interests.“Google can brainstorm for you”Reid points out.
But Gemini goes beyond conversation, reasoning, and planning, which already represents a radical advance. The next step is the greatest possible humanization, giving it not only hearing but also another fundamental sense: sight. Demis Hassabis, director of DeepMid, He explains: “We always wanted to build a universal agent that would be useful in everyday life. That’s why we made Gemini multimodal from the start. Now we’re processing a different flow of sensory information. These agents can see and hear what we’re doing better, understand the context we’re in, and respond quickly in conversation, making the pace and quality of interaction much more natural.”.
Hassabis demonstrates these skills, which will be available in the Live app for Advanced plan subscribers, in a single-shot video recorded in real time. The search engine uses the mobile phone's camera to record the real-world context of a user asking him what he sees, the name of the specific part of an object he's pointing to, how to solve a math problem written on paper, and how to improve a data distribution process in a diagram shown on a whiteboard.
Finally, he asks, “Where did I leave my glasses?” Gemini, who has recorded everything he has seen during the interaction, even if it's not relevant to the conversation so far, reviews the perceived images and answers exactly where he saw them. From then on, the glasses interact with Gemini.
“Gemini is much more than a chatbot. It’s designed to be your personal assistant,” explains Sissie Hsiao, Google’s vice president and CEO of Gemini, referring to the Astra project led by her colleague Hassabis. This is what Sam Altman, head of OpenAI, Google’s competitor and developer of the similar ChapGPT-40, calls a “super-competent colleague.”.
“The responses are personalized [you can choose from 10 voices and the system adjusts to the user’s speech pattern] and intuitive to maintain a real back-and-forth conversation with the model. Gemini is able to provide information more succinctly and respond in a more conversational way than, for example, if you are only interacting with text,” Hsiao explains.
There has also been progress in power, not only with new components such as proprietary processors (the Axion chip and the Trillium TPU), but also in charging capacity. Gemini 1.5 Pro subscribers will be able to manage up to one million tokens, which, according to Hsiao, is "the largest window of context." A token is the basic unit of information.
It can be understood as a word, number, symbol, or any other individual element that constitutes part of the program's input or output data. With this capability, Gemini can load and analyze a PDF of up to 1,500 pages or 30,000 lines of code, or an hour-long video, or review and summarize multiple files. Google expects to offer the two million tokens.
To facilitate the implementation of these skills on devices with less capacity, such as mobile phones, Google has updated the specific systems for these terminals and developed Flash, a high-performance system that provides speed, efficiency and lower power consumption.
And although it wasn't the main focus of this year's Google I/O, Google also presented improvements to its artificial intelligence programs for photography, with version 3 of Image, video creation (Veo), and music, with Lyria and Synth ID. The Ask Photos search engine, which will begin operating in the summer, will be able to locate and group images by theme at the user's request and create an album with all related images.
Source: www.elpais.com
News
News
News
News