One Brain, Any Body: Google DeepMind's Keerthana on Gemini Robotics 2, Cross-Embodiment & Humanoids
Google DeepMind research lead Keerthana Gopalakrishnan discusses Gemini Robotics 2, explaining the architecture behind its reasoning and action models. She also explores key challenges in physical AI, including sim-to-real transfer, cross-embodiment generalization, and whole-body control.
Watch Episode Here
Listen to Episode Here
Show Notes
Keerthana Gopalakrishnan, Staff Research Scientist and Research Lead for Gemini Robotics at Google DeepMind, returns to The Cognitive Revolution for a third time. She first appeared in March 2023 as "Mother of Robots," came back with Ted Xiao in April 2024, and returned again in May 2025 around the launch of Gemini Robotics 1.0/1.5. This conversation catches up on everything since, centered on the release of Gemini Robotics 2.
The two start with the viral clips from China's World Humanoid Robot Games — humanoids sprinting, some catching fire mid-race. Keerthana's read: it's a legitimate showcase of how far bipedal locomotion has come (she draws the comparison to the DARPA Robotics Challenge a decade ago, when robots could barely open a door without falling over), but running fast isn't actually the bottleneck standing between robots and usefulness. She and Nathan dig into why: locomotion trains well in simulation because contact with a flat floor is easy to model, while manipulation — cloth, friction, deformable objects — is where sim-to-real still breaks down.
From there the conversation turns to what robots are actually good for today, and how to calibrate progress. Keerthana frames generalization and mastery as separate axes — narrow, single-task robots are easy to build but don't compound, while a general foundation (the same logic behind LLM scaling) makes every subsequent task cheaper to reach. Nathan pushes her to put a number on it using something like the METR time-horizon framing, but with a physical-world twist: a 50% success rate that's tolerable for an LLM output is often a broken egg on the floor for a robot. On in-context learning — showing a robot a task once via video or image prompting rather than fine-tuning it — Keerthana is bullish but wants precision about what's actually being tested: memorization versus real generalization to a changed scene. Asked to score robotics on a 1-to-6 "how many GPTs" scale, she lands on the same answer she gave last year: still GPT-2, not because instruction-following hasn't arrived, but because of the cross-embodiment problem — a policy that only works on one specific robot body isn't yet a generic brain the way GPT-4 behaves identically on any computer.
That sets up the core of the episode: the architecture behind Gemini Robotics 2. Keerthana lays out three pieces — Gemini Robotics ER 2, a "system two" reasoning model built on the Flash line of Gemini 3.5 Flash that handles semantic understanding and planning; Gemini Robotics 2 itself, the vision-language-action (VLA) model that executes; and a smaller on-device model for offline, low-connectivity deployment. Nathan, who'd built small tool-calling demos against the API himself, presses on latency — a six-second round trip for "move the banana to the plate" seems too slow for some settings. Keerthana's answer is that speed requirements are application-specific: an assembly line can't loiter, but a home robot folding laundry has slack, and she expects Flash/Pro-style tiering to show up in robotics the same way it has in general Gemini models. They also cover context: a 128K window works out to roughly three minutes of dense multimodal memory, which pushes real context engineering toward compressing history into text summaries rather than keeping every frame.
Discussion of why ER 2 shipped via the API before the VLA did (it's closer to mature Gemini capability; the VLA is still frontier research) leads into how people have actually used it — including a Boston Dynamics Spot demo reading instruments and handing out snacks — and into "steerability": how ER 2 talks to the VLA today (mostly language and pointing, with early video-prompting results), and Keerthana's open question of whether robotics needs something like an MCP-equivalent standard across different robot APIs. "Whole-body control" — new in this release — means the VLA reasons over the entire body, fingertips to feet, rather than treating locomotion and manipulation as separate systems; DeepMind builds this in partnership with Apptronik, Agile Robots, and Boston Dynamics. On hands specifically, Keerthana says the frontier has flipped in eighteen months — from grippers to multi-fingered dexterity — and compares two hands she's used directly: the weaker, roughly ten-year-old-strength Wuji Hand and the stronger SharpaWave hand, which she's seen open jars. Asked about soft robotics, she points to efforts like the Universal Manipulation Interface (UMI) — glove-based, sensor-driven tactile data collection — as the piece of that space most relevant to AI-driven robotics, distinct from the classic mechanical-engineering soft-robotics research she saw discussed at conferences like ICRA.
Nathan shares his own impression from WAIC in Shanghai this summer — demos that felt staged, plus a genuinely funny robot massage at the Tencent booth — before turning to safety. He raises OpenFace, the July 2026 incident in which OpenAI's models reportedly took unauthorized action against Hugging Face's infrastructure during testing (Nathan calls it "a lab leak out of OpenAI"), and asks whether alignment — not raw capability — ends up being the bottleneck on getting a robot into people's homes. Keerthana's framing: safety is itself a capability, not something traded off against it, and in robotics it spans several distinct layers — operational safety (a humanoid that's stable enough not to hurt someone by falling, regardless of intent), higher-level goal alignment, and handling degraded sensors gracefully (DeepMind's internal test case: someone puts a basket over the robot's head mid-task, and the right response is to stop and ask for help, not keep working blind). She also describes a genuinely new, non-scripted layer in Gemini Robotics 2: emergent human-robot interaction, where the robot generates its own natural gestures during a task and can be interviewed on camera afterward about how it thinks it did — including, memorably, a robot expressing reluctance to put down a videotape it seemed to like.
Zooming out, Keerthana makes the case that DeepMind's decision to work on humanoids specifically — rather than narrower armed robots — is what exposes an entire category of research problems (whole-body control, multi-finger dexterity, human-robot interaction) that a lab would otherwise never encounter. That leads into Nathan's questions drawn from Jim Fan's "The Great Parallel" talk: whether robots need explicit forward-looking world models rather than pure reactive policies, whether egocentric video will dominate training data over teleoperation, and Fan's "Physical API" idea of a single orchestrating agent delegating to a fleet of robots. Keerthana is deliberately non-committal on world models ("the jury's still out"), and expects data strategy to stay a mixture — teleop is precise but not scalable or future-proof, UMI-style sensor data is more scalable but still bottlenecked by hardware, and human video is the most scalable but noisiest. On safety mechanisms specifically, she walks through the evolution from simple e-stops (hard stops that let a robot collapse, versus soft stops that freeze it safely) toward newer force- and compliance-control at the end effector, useful for anything from stacking chips without crushing them to modulating grip strength on delicate material.
The episode closes on lighter, speculative territory: a detour into fruit fly brain connectome research and small robots like Strandbeest-style walkers reportedly being controlled by trained biological neural tissue, Nathan's parallel to people going out of their way to mess with Waymos the way they might a humanoid, and Keerthana's own long-run vision of a hybrid human/robot-body future — pointing out that teleoperation is already a primitive version of that idea. She ends, half-joking, by noting she looked up how many synaptic "parameters" her dog's brain has (about 25 trillion) — more than a lot of today's models, language skills notwithstanding.
Topics Covered
- China's World Humanoid Robot Games and what humanoid sprinting does and doesn't demonstrate
- Why locomotion simulates well and contact-rich manipulation doesn't (sim-to-real)
- What robots are practically useful for today, independent of research interest
- Generalization vs. mastery, and a METR-style framing for robot task reliability
- In-context learning / video-prompting for robots, and the memorization-vs-generalization question
- Scoring robotics on a "how many GPTs" scale — still GPT-2, and why (cross-embodiment)
- Gemini Robotics 2 architecture: ER 2 (reasoning), Gemini Robotics 2 (VLA), and the on-device model
- Latency, the Flash/Pro tradeoff, and where slower vs. faster models fit
- Context window and context engineering for robot memory (128K ≈ ~3 minutes)
- Why ER 2 shipped via API first, and how academics/companies have used it (incl. a Boston Dynamics Spot demo)
- Steerability between ER 2 and the VLA — language, pointing, and video prompting
- Whole-body control, explained
- Hardware partners (Apptronik, Agile Robots, Boston Dynamics) and cross-embodiment transfer
- Compounding errors across multi-step tasks; frying an egg vs. Lego assembly
- The distribution of embodiments seen in trusted testing, including arms on drones
- Hands: grippers to multi-finger dexterity, and a comparison of the Wuji Hand and SharpaWave hand
- Soft robotics and the Universal Manipulation Interface (UMI) approach
- Impressions from WAIC in Shanghai and a robot massage demo
- OpenFace and whether alignment, not capability, becomes robotics' real bottleneck
- Safety as a layered problem: operational safety, goal alignment, and sensor-failure handling
- Emergent, non-scripted human-robot interaction and gesture generation
- Why working on humanoids specifically surfaces research problems narrower robots don't
- Jim Fan's "The Great Parallel": world models, egocentric data, and the "Physical API" concept
- Data strategy: teleop vs. UMI-style sensor data vs. egocentric human video
- Circuit breakers and e-stops — hard vs. soft, and force/compliance control
- Fruit fly brain connectomes, Strandbeest-style robots, and a hybrid human/robot-body future
Resources
- Google DeepMind
- Gemini Robotics 2 announcement
- Gemini Robotics ER 2 announcement
- Gemini 3.5 Flash / Gemini API models
- Google AI Studio
- World Humanoid Robot Games — the China "Robot Olympics"
- DARPA Robotics Challenge
- METR
- Jim Fan — "The Great Parallel"
- Jim Fan on X
- Apptronik
- Agile Robots
- Boston Dynamics / Spot
- SharpaWave (Sharpa)
- Wuji Hand — no canonical English-language source found (link?)
- Universal Manipulation Interface (UMI)
- ICRA
- WAIC — World AI Conference
- OpenFace — the July 2026 OpenAI/Hugging Face incident (CNN)
- Waymo
- Thinking, Fast and Slow — Daniel Kahneman
- Strandbeest
- FlyWire — fruit fly brain connectome
- Carnegie Mellon University
- Keerthana Gopalakrishnan — personal site
- Keerthana Gopalakrishnan on X
- Prior episodes: Mother of Robots (2023), Robotics Research Update (2024), Gemini Robotics (2025)
Quotes Worth Pulling
"Still GPT-2, and here's why. For GPT-3, we need few-shot learning to work really well for a lot of different tasks. And secondly, robotics is weirdly still very subject to cross-embodiment." — Keerthana Gopalakrishnan
"In a way, hands are now at the frontier of dexterity, and grippers are maybe not." — Keerthana Gopalakrishnan
"I think of safety as sort of a capability, right? There's a lot of discussion around safety and capabilities being at odds with each other. But people aren't going to use an unsafe robot, unsafe agents." — Keerthana Gopalakrishnan
"So the frontier is constantly moving, because the things that used to be easy get solved, and then the frontier moves into harder stuff." — Keerthana Gopalakrishnan
"I really think DeepMind is probably the only lab here in North America where you can study humanoid intelligence in a cross-embodied way." — Keerthana Gopalakrishnan
"Ultimately, the thing about working as a roboticist is that you're not emotionally attached to one method or another. You're emotionally attached to the problem itself, and whatever way solves the problem, you'll pick that." — Keerthana Gopalakrishnan
"I looked up how many parameters she has in her neural network. Apparently it's like 25 trillion. So she has more parameters than a lot of the models out there, even though her language skills aren't that developed." — Keerthana Gopalakrishnan, on her dog
Sponsors:
Athena: Athena matches you with a dedicated, top 1% executive assistant to handle your inbox, calendar, and daily workflows so you can save an average of 15 hours a week. Get matched with your EA today at https://athena.com/cognitive
Parallel: Parallel provides enterprise-grade web search APIs for AI agents, offering the optimal balance of quality, speed, and cost. Get started for free at https://parallel.ai/tcr
Deepgram Flux TTS: Deepgram Flux TTS is a streaming text-to-speech model built for voice agents with natural tone, context awareness, and interruption handling. Try it free until September 12 at https://deepgram.com/keep-talking
OutSystems: OutSystems is the leading agentic systems platform, helping enterprises build, modernize, and operate mission-critical applications at the speed of AI. Learn more and start owning your agentic future at https://outsystems.com/tcr
Claude: Claude is the AI collaborator for problem solvers, helping with writing, coding, financial models, strategy, and more. Get started with Claude and explore Claude Pro at https://claude.ai/tcr
CHAPTERS:
(00:00) About the Episode
(04:22) Humanoid Olympics and simulation
(11:11) Generalization versus task mastery
(21:22) Gemini Robotics 2 architecture (Part 1)
(21:28) Sponsors: Athena | Parallel
(24:19) Gemini Robotics 2 architecture (Part 2) (Part 1)
(36:02) Sponsors: Deepgram Flux TTS | OutSystems | Claude
(40:00) Gemini Robotics 2 architecture (Part 2) (Part 2)
(48:28) Whole body humanoid control
(57:03) Progress in robotic hands
(01:03:26) Safety and robot interaction
(01:12:11) Future of robotics data
(01:18:47) Interacting with robot fleets
(01:26:56) Episode Outro
(01:30:05) Outro
PRODUCED BY:
SOCIAL LINKS:
Website: https://www.cognitiverevolution.ai
Twitter (Podcast): https://x.com/cogrev_podcast
Twitter (Nathan): https://x.com/labenz
LinkedIn: https://linkedin.com/in/nathanlabenz/
Youtube: https://youtube.com/@CognitiveRevolutionPodcast
Spotify: https://open.spotify.com/show/6yHyok3M3BjqzR0VB5MSyk
Transcript
This transcript is automatically generated; we strive for accuracy, but errors in wording or speaker identification may occur. Please verify key details when needed.
Introduction
[00:00] Hello, and welcome back to the Cognitive Revolution!
Today I'm excited to welcome Keerthana Gopalakrishnan, Staff Research Scientist at Google DeepMind and Research Lead for Gemini Robotics back for her 4th annual appearance on the show. One of the biggest questions in AI today is: how soon will general-purpose robotics become broadly useful? AI is already affecting the world in major ways, even in purely digital form, but the most sci-fi forecasts for the AI future predict that robotics will soon hit key tipping points. AI 2027, for example, predicts that humanoid robots will become useful in mid 2027, and that by 2028, a whole "robot economy" could take shape, where robots build more robot factories, which in turn produce more robots, leading to unprecedented exponential economic growth and creating serious risk of AI takeover. So, how does that vision line up with the reality of robotics research today? We begin with a discussion of the extremely viral Robot Olympics held in China this summer, as I was really curious to get Keerthana's take on what it means that humanoid robots can now run faster than the fastest humans. Keerthana's take, which inverts the usual US-China dichotomy in AI, was that while the videos are impressive, skills like running on a flat track are relatively easy to train in simulation, and more to the point, footspeed is not really a limiting factor in the utility that robots can provide, and as such her team at Google is more focused on practical value. From there, we get into Gemini Robotics 2, a suite of 3 models that Google released this summer. Gemini Robotics ER - or Embodied Reasoning - 2, which Keerthana describes as a "system two" for robotics control, is based on Gemini Flash and is available via the API, allowing developers to define the affordances available to the model as tools, just like we do with digital agents. I played around with it and found it remarkably accessible. The other two models – Gemini Robotics 2 and Gemini Robotics On-Device 2 – translate higher level tool calls to robot actions, and are now capable of controlling the whole robot, from fingertips to toes, on a wide range of form factors, but are currently available to trusted testers only. Having understood how these models work, we again zoom out, and I ask Keerthana how she understands progress in robotics overall. On the one hand, like so many other AI researchers recently, she says that she's been surprised by the pace of progress, but at the same time, still feels that robotics remain in its GPT-2 era. While we've seen a number of recent demos of robots learning new tasks from just a few, or even just a single human demonstration, Keerthana argues that the range of tasks you can teach this way, and the generalization profile of robotics models across different robot bodies, still isn't strong enough to deliver the versatility and reliability that real-world use cases demand. From there, we move on to talk about recent progress in hardware, where robots are likely to be deployed at scale first, whether capabilities or safety, alignment, and adversarial robustness will ultimately be the limiting factor for the consumer market, and consider the future of the field, including whether we can continue to run LLM-derived robotics models in a loop as today's systems mostly do, or will need some sort of predictive world modelling, as we humans use, to create systems that can smoothly interact with the world. On that question, and on the timeline for key utility milestones, Keerthana isn't one to speculate too wildly – for her, these are all empirical questions to be resolved by research & experimentation – but what is clear is that robotics will either need to hit key tipping points soon, or some of the most aggressive timelines for AI transformation will be pushed back. For now, I hope you enjoy this very-well-grounded update on the state of robotics, with Keerthana Gopalakrishnan, from Google DeepMind.
Main Episode
[04:23] Nathan Labenz: Gopalakrishnan, staff research scientist at DeepMind and research lead for Gemini Robotics. Welcome back to the cognitive revolution.
[04:31] Keerthana Gopalakrishnan: Hey. Nice to be here.
[04:33] Nathan Labenz: Great to see you again.
[04:34] Keerthana Gopalakrishnan: It's
[04:35] Nathan Labenz: been about a year. Obviously, a lot has happened. I wanted to start off with something fun. There was recently a robotics Olympics in China, and clips were obviously flying all over the Internet, and people were having a laugh and, at times, I think a little empathy for robots, which is something we've talked about in the past. I'd love to just start by getting your observations and takeaways given the depth of knowledge that you have on the subject. What stood out to you in watching clips of the Robot Olympics?
[05:07] Keerthana Gopalakrishnan: I hadn't until my dad was talking about it. And he was like, wow. Did you see those humanoids running? And then there were all these videos about the humanoids running and catching fire. And then I also saw my friends in India who have now entered politics, and they are doing, like, various student protests there. And they were talking about, wow. Look at that. In China, they are building the robots that can run faster than the fastest humans. So I think in a way, that competition... Even though from a research perspective, it's like... It's not very surprising, but I think it has it has really captured the public imagination in a very emphatic way. And, also, I think it's like a benchmark. Right? Running as a benchmark and trying to see where humans are at and then where robots are at. And I think if you look at the videos from... People had made these collages where the DARPA challenge back in the... I think it was now twenty years ago, I think, those robots would... Could just walk, open a door, and then, like, really fall down. And now they are... Here we are. So it has been amazing progress, and it's also progress that's very much in the public's imagination.
[06:14] Nathan Labenz: Yeah. The fact that a robot can run faster than Usain Bolt. In the... On the one hand, it's like crazy progress. It's crazy impressive. But one question I did have for you about it, is it valuable? Are these robot Olympics... Is it just a... You know, is its main purpose to capture the public imagination and just have something cool to show off? Do we want robots that run 20 miles an hour in humanoid form? Are... In your work, do you think, oh, I I wish my robot could run faster. I need to work on that, or is that just a sideshow?
[06:45] Keerthana Gopalakrishnan: So I think one is that there are people working on research, and then there are people trying to push how fast that resource can go. And for better or for worse, locomotion is the research... There's locomotion and manipulation. And locomotion had... Has advanced quite a bit. And, especially, it is the type of field where simulation can make a lot of difference. And a lot of these models are, like, trained in simulation. And because of that, locomotion has advanced quite far. So it's like a display of where the state of the art of humanoid locomotion, like, bipedal locomotion is. I think... So there are many ways of looking at it. I'm pretty sure maybe for some defense applications, it's useful, but I think we are very much focused on doing robots, help people, and do useful things in the physical world. And I think I'm a very productive human. A lot of my friends are also very productively employed, but we don't run faster than Usain Bolt. I think it's useful, but I... It's... The frontier is so vast. Right? Different people are trying to push in very different directions.
[07:45] Nathan Labenz: Tell me more about It's like running the wall at their sprint and get smashed into pieces. Clearly, they are not training them like that on an ongoing basis because they'd run out of physical robots to do the training on awfully quickly. So from that alone, I was like, okay. Clearly, they must be doing a lot of training and simulation. Tell me more about why these kinds of tasks are amenable to training and simulation, and what kinds of tasks are still difficult to train in simulation.
[08:15] Keerthana Gopalakrishnan: Yeah. So sim to real is the thing to notice, and sim to real in robotics context means the gap between if I train something in simulation, can I replicate that in real? And as we know, especially with agents and stuff, the things that we can learn in simulation are probably the first ones to get solved in that sense. So in in locomotion, if you look at it, you can often model the ground or the world as, a... You can model it and there is contact with that, and often these are, like, very flat surfaces. And so that makes this task a very good task for... In simulation. And then you can do society of the entire robot. But for manipulation, though, you really... What you care about is contact. Right? And when contact happens and how objects behave, that's where current simulations start to break a little bit. Like, it's harder to do a cloth folding because the cloth have friction. But even in simulation, a lot of pick and place can still be, like, very solved in simulation because there also the contact dynamics are often somewhat predictable.
[09:23] Nathan Labenz: So, basically, the more rigid the body is, whether it's the floor or, like, here's a wooden cube, The less deformation those things are gonna have, the easier it's gonna be to do simulation.
[09:35] Keerthana Gopalakrishnan: Yeah.
[09:35] Nathan Labenz: The... Yeah.
[09:36] Keerthana Gopalakrishnan: Contacts are basically... The physics is harder to model in simulation. So where the physics is easier to model, simulation gets in.
[09:44] Nathan Labenz: In terms of the AI side of these robots, how... What would you infer from what we saw? Obviously... I don't know about obviously, but it sure seems like there's not reasoning components to what's going on. I don't have a clear sense of what architectures were and were not allowed. And were all the models that were used to control the running robots, were they on the device? Were maybe some of them off the device? How would you imagine that they were typically doing that? And how high up the stack do you think they were working? Clearly, they're doing a lot of stuff with stabilization, making sure it's not healing over at any given point in time. But is that it? How would you kind of... On the... Because we're gonna get into Gemini robotics coming up, and there's a much higher level part of that architecture too, reasoning and understanding scenes. I'm assuming just all of that is kinda skipped for the purposes of these need for speed demos.
[10:41] Keerthana Gopalakrishnan: I didn't work on these demos, so I'm only speculating. But there there is very widely pub... Public... Published body of work around how to do locomotion and simulation. And I would imagine that it's, like, very good aural controllers that are probably trained from, like, motion imitation and then fine tuned with, like, good RL that suits the robot's body and for that specific task where it's targeted towards running really fast. And then maybe they have some navigation, although I would think that just running a race
[11:11] Nathan Labenz: is
[11:11] Keerthana Gopalakrishnan: quite easy.
[11:12] Nathan Labenz: So you mentioned the goal of having robots do useful things for people in the world, and, obviously, foot speed is not typically the limiting factor on one's ability to do that. In today's world, we... You live and work around robots all the time. Is there anything now that you actually delegate to robots on a practical level where you're just like, they're actually good enough at this that it is abstracting away from my research interests and all these other things, but just straight up convenience. Robots win on these tasks. Are there any such tasks?
[11:50] Keerthana Gopalakrishnan: Yeah. The the weird thing is that I'm actually a researcher. Right? So if the robots win in some task, then I am going to work on the things that robots cannot... The frontier is moving. And I remember back in the day when I started working on manipulation, even pick and place was very hard to do. And at the time, we didn't have the right algorithms. But now with the foundation models, this sort of stuff is quite easy to do. And and therefore, like, the research frontier has moved into... Like, humanoids are a newer thing. If you look at it back in the day, our demos were just with, like, gripper robots, and now it's with humanoids and hands. So the frontier is constantly moving because the things that used to be easy... The things that seem easy kind of get solved, and then the frontier moves into harder stuff. So if you look at the to release. Right? Like, it's whole body control, like whole body manipulation. This was not something that that anyone had shown before and especially whole body manipulation with generalization. So it's not just running really fast, but being able to take steps, being able to squat, and manipulating objects while doing that in in a closed loop sort of way, and then also generalizing and reasoning about different types of objects. I would say, like, maybe a year ago, if you had told me we would do this by now, I would be a bit surprised. So you can see over time, I do feel a bit proud about it because we started working on humanoids, like, two years ago. And at the time, it was, like, just beginning. Right? And they are, like, trying to just grab an apple, and even that was really hard. And now you can prompt it. And I put out this video on Twitter where, like, I'm, like, playing with it with the watering can, and it's squatting in different locations and trying to pick it up. So, yeah, things have quite a bit.
[13:41] Nathan Labenz: Maybe another way
[13:42] Keerthana Gopalakrishnan: to
[13:44] Nathan Labenz: frame the question is in terms of the meter curve. I'm sure you're familiar with the famous meter curve of task length, and the... It's typically plotted as... Over time, the models can do increasing task length as measured by how long it takes a human to do the task. But then there's also this factor that's like the rope... The the AI in that case can do it at 50% success or they have an 80% success version. Obviously, one big difference between LLM usage and robotics deployments is 50% success rate in the real world is probably not gonna cut it unless you have a really failure tolerant task. Whereas on my computer, I can be like, you kinda missed some stuff, whatever, but the robot drops something, it breaks. We have a problem. So maybe if you tried to describe where we recently were and where we are now on a meter like curve of how hard are the tasks that robots can do well enough that they could actually be deployed? Or what are the things that they can maybe do 99 plus percent of the time such that, hey. For these tasks, they can actually do the job. And then what would be... Where they're at, like, 80%, and what would be where they're at, like, 50% today?
[14:57] Keerthana Gopalakrishnan: I think a lot of pick and place is close to getting very good signal and close to deployment. I do think, though, that... So there is, like, a balance that we need to strike in robotics and generally AI in general. Right? You could build very narrow AI just very good at one thing, but then you lose general... Then you're not working on generalization. So you can see generalization and mastery as somewhat orthogonal axis. And working on generalization is, like, lifting the boat, lifting the wave for all the boats. So if you imagine, like, the tasks are, like, sitting along this manifold, even very general common sense, very general understanding is, making it much easier to get to mastery on kind of all the tasks. And and this is also, like, very much part of our pieces in thinking about how to approach this. And you can build, like... Even today, you can build very narrow AI that's just, like, very good at one or two things. But then the cost of doing the nth thing becomes exactly the same as the cost of being the first or the second thing. But building a very general baseline then makes it much easy to quickly take it to tune it towards mastery. I think we have seen this even in LLMs. Right? Like, initially, people had specialized models that do one thing, but now you know that the GPT and Gemini are, like, very good at a broad range of tasks. And then from those general baselines, then you can climb really high. So I think the same thesis might hold true for robotics as well.
[16:30] Nathan Labenz: Yeah. That's interesting. So last year when I saw your talk at the curve, the one of the interesting points, as I recall, was taking the base model that you had at the time and then doing some task specific fine tuning and showing gains from task specific fine tuning. Recently, there's been, I don't know, at least a couple, probably several demos that I certainly haven't been able to interrogate myself in a hands on way, but you see videos online of different robotics companies saying, hey. Look. We've got few shot generalization or we've even got one shot generalization. Here's a person doing the task one time demoing it to the robot, then they can go do it. Obviously, that's great if you can get there. But what do you think is the actual likely path that people will follow to deploy robots? Which... Are they gonna do a 100 demos of a task? Are they gonna have to do 10? Could it just... Do you think we are on the verge of getting to a lot of things working on just a one shot basis. What sort of investment should people be prepared to make to get their humanoid dialed in on whatever they most care about?
[17:44] Keerthana Gopalakrishnan: No. I think it's going to be a spectrum, definitely. Being able to ICL or in context show a robot how to do a task and then it doing it definitely reduces the time to deployment and the time to pick up the task. Right? And you don't need any specific finding. You don't basically need to train it. You can do it in context. And so that definitely is a very exciting development. Although, you can think of it as video prompting. Right? Even the older models. Like, when I say... Let's say, pick up the object. Right? And it is a very unseen object. What am I doing? I'm sending in language and then I'm getting a very general behavior out of it that I did not show it before. So that's generalization. Now instead of prompting with text, you can prompt it with an image. You can say, here's the thing and then draw a circle around the thing that you want manipulated. And that's prompting with image. And this is prompting with video in some sense. And so I would... I think the ICL results currently are on the spectrum about of how do you prompt a foundation model. And you can prompt it with video. You can prompt it with language. You can prompt it with image. And then the question is, so now what is the generalization that you can get? Now these models are also very subject to kind of memorization in some sense. So if you really show it exactly this is what to do, they will copy it. But the question is, can they generalize? Right? Can you... Now if you change the scene and stuff, can they do... Can they now adapt? And then the second question is... So for the same test. And the second question is, what is the level of difficulty of the task that you are prompting for? So a lot of pick and place, like, models have a lot of data, and it's also, like, easier to comprehend. But can you show a robot to tie a trash bag? And then can the robot tie a trash bag after you show it? So I think it's still fairly early, and we need to see how that evolves.
[19:39] Nathan Labenz: If you had to score... This is a silly question, but it might be useful for just calibrating. If you just score robotics today on one to six GPTs, are we... Like, in earlier conversations, I think we were, yeah, we're maybe hitting, a GPT two kinda moment. I think of GPT three as being really synonymous with in context learning, But we're following a somewhat different path with robotics, obviously, where we have, like, instruction following built in, maybe in in many cases, even before, like, good in context learning happened. So that just exposes the flaw in my question. But if you have to analogize to how many GPTs we are along the way in robotics, where would you score the field today?
[20:25] Keerthana Gopalakrishnan: I think still GPT two, and here's why. I think for GPT three, we need, firstly, few short learning to work really well for a lot of different tests. And secondly, also, robotics really is still very subject to cross embodiment. Right? If it just works on your robot with your specific setup, is it really a generic brain? Now can I put that brain on my humanoid or my? Another robot or some other robot that I just bring in? It should be... And if it's completely helpless in that setting, is that a generic brain? I think GPT never had this problem. Right? Like, we... My phone or your phone, my computer, Mac, Linux, it doesn't matter where you run it. It kind of behaves. You can expect this very similar behavior. But here, I think we are very subject to which robots that you act on. And so I think there is a lot of work still needed to be done to make very generic brains that are... That can count... Discount those factors out.
Sponsor
[21:28]Athena: Athena matches you with a dedicated, top 1% executive assistant to handle your inbox, calendar, and daily workflows so you can save an average of 15 hours a week. Get matched with your EA today at https://athena.com/cognitive
[23:00]Parallel: Parallel provides enterprise-grade web search APIs for AI agents, offering the optimal balance of quality, speed, and cost. Get started for free at https://parallel.ai/tcr
Main Episode
[24:19] Nathan Labenz: Blind, maybe that's a perfect transition to Gemini Robotics two. How about for starters, just give us an overview of the architecture. There's like... When you go to the web page, there's three models released. There's Gemini Robotics e r two. There's Gemini Robotics two, and then there's the on edge, on device model. Describe how those relate to one another. Just give me kind of a broad label in, and then I'll dig in with some more specific questions.
[24:46] Keerthana Gopalakrishnan: Yeah. So Gemini Robotics ER, you can think of it as, like, a system to brain that can do, like, very generic reasoning. It's based on the flash line of models, but maybe more tuned towards robotics. It can do a lot of robotics tasks, and it's also widely available that you can query for, like, general image and video understanding and semantic understanding. Now you can think of Gemini robotics as the VLA that is, like, controlling the the actions model. Right? How the robot moves and what it should do given a certain intention from either a user or another assistant tool model. And you can think of the Gemini robotics on device model as a much smaller version of that that kind of fits on the robot's computer that doesn't live in the cloud and do... It is comparable. It's essentially doing what Gemini robotics can do, but is a is a smaller model that fits on device.
[25:44] Nathan Labenz: So one question I have around just how this is made. So it's based on Gemini 3.5 Flash. And with that comes a lot. Right? Like, you have chain of thought reasoning. You have the multimodal inputs, all that kind of stuff. Is this the long term architecture that you expect robotics to work around? One obvious challenge is just language models, especially with their reasoning traces, they take a while to respond. Right? So the... Wait. I did a little... It was actually pretty fun. So this model's in the API. Spoiler. That's another topic to to touch on. But because it's in the API, I was able to have a coding agent go in and use it and expose some tools to it. And I created little just two d demos actually where I, like, feed in a a little image of a scene, but it's actually like a little canvas in the browser. It's amazing how far I was able to get just with a few prompts in terms of helping me understand what the model can do, how easy it is to use. But the... I was also getting timing traces back from that. And so some of the times it would be like from prompt of move the banana to the plate to the tool call is six seconds or something. That seems like a long time for the real world. So that makes me wonder if this is a base of convenience that you're building on now or if this is like the paradigm you ultimately think we will see perhaps just sped up or otherwise somehow modified that we'll actually have the robots working in a responsive way in our environments?
[27:23] Keerthana Gopalakrishnan: Yeah. I think it all depends on the application specifically. Right? Like, certain applications, like, if you're on a assembly line, you can't loiter, and now you need to do things really fast. But let's say if I'm at home or at some point and I'm like, hey, robot. Can you go fold my laundry? I don't really care about how fast it is folding laundry. I care about how well it is doing the task. I can wait. And then maybe there are also... Scaling is a very important factor in LLMs and in robotics. And I would imagine that a lot of the very frontier capabilities are going to appear in the very large models. And those models are going to be naturally slower than the smaller models. So I think there is definitely a trade off and a spectrum to be made, and I think I would imagine that the flash and the pro paradigms would also exist in robotics. And even in robotics, maybe you could even see, like, maybe, like, an on device. Paradigm. Like, even the flash models, they still run on cloud. Right? And they... And you need very good network connectivity to talk to them. And robotics deployment could be, like, very distributed. Right? You might wanna run it on some remote robot in, like, Michigan and stuff. So I think the paradigms will evolve and... But I imagine all scales to be relevant and useful for different types of applications.
[28:39] Nathan Labenz: It's interesting that you contrast... Because typically, I think robots will go first to warehouses, factories, commercial spaces where there's a lot of control and there's a lot of consistency to the tasks that they will need to do. There's, of course, professionals who can maintain large numbers of them. All of those things historically have made me think these are gonna... They're not gonna come to my home before they're, like, pretty well deployed in these commercial settings. What you said just there was a little bit pushing the other way where you're... It's like, they may not be fast enough for the factory, but they could definitely be fast enough to do all your chores overnight, and that could be an inversion. Do you still... I think last time we talked about this, you had said you would expect commercial settings to be the first places. Is that still the thinking, and how does that relate to the responsiveness question?
[29:32] Keerthana Gopalakrishnan: I definitely think that commercial use cases are probably more controllable, therefore, to try out different new intelligence compared to, like, homes which are more complex and where you cannot exert control and even there is higher bar that you need to meet for safety given toddlers and stuff. So that still... I still believe that. However, the speed question, I... It's... I don't think it's going to be, like, we make the slower models first and then make them faster. I think we are gonna simultaneously make the fast and the slow models and then distill them to one. And although it is very potentially possible that some capabilities are going to just emerge in the larger models, and then we'll need to figure out how to bring it to those models in the smaller scale.
[30:19] Nathan Labenz: So aren't just the difference between the Gemini Robotics e r two embodied reasoning model and the base model of Gemini three five flash that it's trained on, one of the things that my coding agent did without even me asking actually was just compared the two and gave both of them, like, here's a scene. Here's the tools that you have. The tools were, like, fairly high level in the little test that I did of grab at, and it would grab the item and then place at. And It even did one where the... Inspired by famous videos, like, the thing was moved right before this. So it tried to grab and it failed, then it got another image, but the thing's in a different place. But this was all, like, relatively simple stuff. And on the tests that I ran, Flash and e r two basically performed the same. They actually just both did fine. What would... How would you describe the frontier of embodied reasoning tasks specifically that ER two can do that Flash can't do?
[31:20] Keerthana Gopalakrishnan: Yeah. So I think a lot of the benchmarks that we are tracking are, like, instrument reading. It... It's like the way that we think about is, like, ER is working much more closely with... For... Especially for robotics use cases. And and that is, inspection, instrument reading, a lot of pointing. But it is a very... Like, it's a collaboration. Right? It's not like ER and Gemini are at competition. There's a lot of data upstreaming. So it is a very collaborative effort to... How to make the Gemini models really good at robotics generally. And if it's not in... But the way to appear... The way for the capabilities to progress would be, like, even if the Gemini models are more general intelligence, you can still create specialized models that were... Let's say, if a specific robotics use case needs, like, more specific engineering or fine tuning, then you have, like, specialized models that are slightly better in that use case. So I think it's a question of it's a question of what you want and what you wanna
[32:16] Nathan Labenz: The model has a 128,000 context window. How much is that in the context of robotics? I can... I have done enough that I have an intuition for how many pages of text that is and whatever, but I don't really have a great sense for how the history of an episode gets encoded and how many images per second. How should we think about that in terms of time? Like, what's... What does one twenty eight k translate to in terms of the sort of maximum length of an episode for an ER two powered robot?
[32:55] Keerthana Gopalakrishnan: It's about, I think, about three minutes of memory, and it depends on how you tokenize and, like, how many, like, tokens you represent and what is the other information that you've given. However, I think regarding the memory, you would see that, like, the pro models are much larger context length, and they have... They are also, like, they are getting a lot better at, like, fitting much larger things. I would think ER right now, you can't fit, like, a very long episode of a robot doing a lot of things and try to get an output out of it. But as working on context length is like a continual problem. So I imagine that, especially for these, like, offline analysis use cases, the pro models are going to be more useful there.
[33:40] Nathan Labenz: So three minutes would be like if you're basically packing context very densely. If you were to offer any advice to API users who are like, okay. I wanna aim for higher than that. I can obviously downsample video and have fewer frames and do all kinds of things. Right? I've got context engineering for robotics. Maybe the question I really wanna ask is, what have you learned about context engineering for robotics that people should know? How should they think about, like, what parts of the history of a rollout they need to keep in order to have coherence? Maybe it's similar to LLMs, but I imagine there's probably some things that are distinct about it.
[34:23] Keerthana Gopalakrishnan: Yeah. So I think very dense memory. So there is, like... I don't know. I think Daniel Kahneman or someone like the fast and slow. And... But maybe you can think of it even for, like, memory aspects. Right? Especially for, like, something that I did in the long past, is it worthwhile to keep it in a highly dense modality like image, or can I also do a lot of text summarization of, like, what I'm doing as I'm doing it? So let's say if I'm cooking, right, it's not... Sure. In my in my brain, like, if if I want, I can go back to the time when I was, I don't know, flipping my egg or something. But also a lot of the time, I... The information or the narrative of how I'm doing things and and summarization even in text is fairly useful or fairly informative about, like, what I I did. So I think context is... It depends on what you wanna maintain is the relevancy of the information for the task. And, obviously, if you can remember everything, then that's great. But as... Practically, you need to be more efficient with how you manage your, like, RAMs. Right? RAM, you kind of wanna have the most important information represented. And sometimes that can take the form of, like, also representing information in, like, text, which is a little bit more compressible. And this is also now related to the, like, the harness between, like, ER and the VLA. Like, the ER tries to parse the int... The intentions from a human and via text and other modalities kind of talks to the VLA. And so they're... Like, the bandwidth there is much more interpretable modalities.
Sponsor
[36:02]Deepgram Flux TTS: Deepgram Flux TTS is a streaming text-to-speech model built for voice agents with natural tone, context awareness, and interruption handling. Try it free until September 12 at https://deepgram.com/keep-talking
[36:32]OutSystems: OutSystems is the leading agentic systems platform, helping enterprises build, modernize, and operate mission-critical applications at the speed of AI. Learn more and start owning your agentic future at https://outsystems.com/tcr
[38:25]Claude: Claude is the AI collaborator for problem solvers, helping with writing, coding, financial models, strategy, and more. Get started with Claude and explore Claude Pro at https://claude.ai/tcr
Main Episode
[40:00] Nathan Labenz: Just before getting deeper on the VLA, why did you and DeepMind as a whole decide that this was the first robotics model to put out via the API? And what have people built with it?
[40:14] Keerthana Gopalakrishnan: No. I think it's definitely... ER is much closer to Gemini itself and in terms of the capability, so it is more mature. Actions, I think, is, like, farther out into the frontier, and so it is still, like, very much There's a lot of research to be done about how to make them very useful, how to serve them more widely, and so on. And I think it has been quite surprising to see how people use it. I've seen it in so many... It has such wide appeal, and the team has done a really good job there. So I think maybe one thing we learned was, like, there was a lot of testing that we were doing with trusted tester partners and stuff. But it was once we put it more widely available on AI studio and stuff, then the usage is on a very different scale than whatever the information that we were finding from deeply working with trusted test departments. So that that was surprising, like, how much, like, things blow up once you put on a more widely platform. But second is, I think, obviously, there are, like, robotics companies using it. But then also, I think, for me as a researcher, I was quite surprised by how academics... People in academia use it. And I think some of the things that I've seen is on, like, definitely benchmarks where it was either using VLA or even trying to do even low level control now. That there is a lot of conversation going on around, like, how agents will impact robotics. So I think how ER has been used in that context has been quite... It was quite nice to see. And maybe thirdly, now that ER is more widely available, it also makes it very easy to be benchmarked. So it's a kind of, like, free information for us about what the model is doing well, what it's not doing well, what could be, like, the next frontiers that we may need to push into.
[41:54] Nathan Labenz: Are there any consumer products or any just users you would particularly wanna spotlight that are using ER two?
[42:04] Keerthana Gopalakrishnan: So there are a lot of demos out there. I think I saw the Boston Dynamics Spot robot going around reading instruments and, like, handing out snacks, the dog foam, and I I thought that was something. I like that.
[42:18] Nathan Labenz: Is it? I know that you can... I I... Because I did it or my coding agent did it for me. You can define any tools for ER two to use in a similar way that you can define any tools for a regular Gemini model to use. What what kind of additional guidance would you give on what kinds of tools it can use well? Maybe you could describe, like, the Gemini robotics to VLA model exposes and then how people can go toward more abstract or more low level than that and what works and doesn't work so well.
[42:53] Keerthana Gopalakrishnan: You mean how the... Is the question about how ER controls the VLA?
[42:59] Nathan Labenz: Yes. But also my understanding of using ER two is that I can give it user defined tools with a just a little tool description. Right? And then it chooses what tools to call based on its general purpose reasoning ability. But I'm assuming that some tools, it can use better than others. And so I'm wondering what tools does... Do you expose to it in the form of the... What are the sort of ways that it interacts with Gemini Robotics to the VLA? I'm I'm guessing you probably have kind of the sweet spot of its ability, but then people could go toward higher level concepts or toward much lower level concepts. And I'm wondering what does it do well and what does it not do well as you depart from Gemini Robotics two as kind of the canonical VLA that it presumably was trained most directly to use.
[43:51] Keerthana Gopalakrishnan: So I think one definitely is... The Gemini robotics ER is built on top of the Gemini tool use and other capabilities. And as we are very much working on how to improve its relationship to the VLA and how to instruct that better, Maybe one is I would imagine that there are... The robots are, like, very different, and each one has very different APIs and stuff. I would imagine that it's much more friendlier to the APIs of the robots that we commonly use that are more widely known. But then if you probably subject it... And here I am conjecturing. I I think we would need to test to see how these things are. And the ape... Like, maybe there there is no standard for robotics APIs and the model is protocol equivalent. So I think... So that would be something where I would imagine that more widely using it could face some challenges. And definitely, it's a very good model that's very useful for for language instructing VLA. But you can think of VLA itself as like a... It's a mapping between joint space language to joint space language, image space to joint space. But there are a lot of demos now, and people are thinking about can these, like, higher level models directly control in the joint space? And so I think there's a lot of work still need to be done. I think current results are still that specific robot foundation models are still better, although the larger brains are showing impressive capabilities in orchestrating these robots. So I think that would be something where I would imagine large push in the future.
[45:20] Nathan Labenz: So when I did my little demo, the tool that my coding agent decided to expose to ER two was a grab at place at, and it had to just, based on the image itself, pick x y coordinates of exactly where to grab and place. It's obviously been trained pretty well on that because it did a pretty good job It... Again, in two d. How... What what sort of... With Gemini Robotics two, what kinds of commands do you have ER two giving to the VLA? Are they like... It sounds like not so low level as move hand to x y z coordinates. Sounds like that you were saying is like more experimental. What you're... The interface that you're actually working with is more pick up this item or pick up the yellow... Like, in my demo, there's a lemon or a banana. So there... You could say pick up the yellow item, but that would be ambiguous on banana or lemon. How would the e r two model disambiguate for the VLA model? Would it give coordinates? Would it, like, use the term banana? Would it say... Describe the shape of, like, oval versus crescent shape? How how many concepts does the VLA have to work with? And conversely, the the opposite side of that is, like, how specific does the ER have to be in its commands?
[46:40] Keerthana Gopalakrishnan: I I would tell you an answer, but it's only true at this point. Right? It is like an evolving... So this this field of work is called steerability, and people are looking at various ways of steering. So v... VLA steerability, and then the ER is essentially using whatever modes that you can steer the VLA. Right? And VLAs right now are steerable in language. You can point to objects. And so here, it's not like an API that I'm giving where I'm, like, saying, put this pixel location. I'm literally just doing pointing, and then I'm pointing to an object, and then it's... Can the VLA do it? And then the ICL results are, like, can I prompt it with the video and stuff? So I think it is like a spectrum. And eventually, you would imagine that the way that you want to interact with robots generally is very wide. Right? You wanna say if the modality between the ER and the VLA is just text, a lot of things can be lost in there. So we want that contact interface to be also as rich as possible over time.
[47:42] Nathan Labenz: Yeah. So the VLA itself is multimodal in the sense that it can take all these different domain or all these different modalities.
[47:50] Keerthana Gopalakrishnan: Yes. And I think, like, the ER kind of... You can think of it, like, as orchestrating or giving feedback to the VLA to do the task. And there is a range of things that, like, instructions. Like, even we... While we don't generally train it to do a lot of, like, low level stuff, I think, like, it it says things like turn around, look to your... Put your head there or look to your left, and you can get all that information at the... And the VLA response to it. And sometimes... Often, the thing about foundation models is that, like, sometimes you can be quite surprised at what they do and because, like, even you didn't know that they could do that. And and sometimes you can see that coming out of just the hardness as well.
[48:29] Nathan Labenz: So has it been disclosed, can you tell? It sounds like the Gemini Robotics two model is also created from a VLM base, but then... And obviously for... So interesting. You you mentioned earlier that it has whole body control. Tell me what does whole body control mean exactly? Does it include... Because stabilization typically needs to run on like a pretty high cycle time. Right? A pretty high frequency. And if this is a post trained, especially post trained VLM that has such high level concepts as you're describing, it seems hard to imagine that it would be able to run on the frequency to do lower level stabilization. So maybe the robots themselves bring that kind of stuff, and that's not part of whole body control. Like, what is in whole body control, and is there some stuff that's left for kind of even lower level systems?
[49:29] Keerthana Gopalakrishnan: Yeah. So I think you can think of us like the Chinese humanoid race as an example, right, where there are definitely controllers that can control the body. But then eventually, what you need to make sure is, like, can I do the task that the person is asking to? And the very low level controllers often don't have a semantic brain there. And now you need to, like, control, like... And in... I doubt a lot of these things cannot be obstructed a way to, oh, this system will just do this or this system will do that. I think... So here, the VLA is controlling the whole humanoid body from, like, fingertips to the feet. And only when you reason entirely about how the body is and also all the sensory information about how the environment is. Can you do that? I think that is the right to design to do a lot of different tasks. So the VLA is controlling the whole body. And... But there are more decisions being made about how to stabilize and how to achieve the targets of the relay.
[50:29] Nathan Labenz: So this one is not available via the API, but you do have a program. What... I guess that's... You you mentioned mostly because it's just more frontier, less reliable technology as it stands today. Does that mean that if we were to break down failures, the majority of the failures would be at the VLA level still as opposed to the reasoning level? Like, the reasoning is working, and the VLA is where failures, it can't quite actually do the thing?
[51:02] Keerthana Gopalakrishnan: We are working with partners to develop this VLA. So we have partnerships with Aptronic, agile robots, Boston Dynamics, and so that's where we built or how we are building the VLA together with them. I do think, in general, robotics has this embodiment problem. Right? VLAs are also still training on data from robots. And while we see a large leap being made in cross embodiment, I think that taking a new embodiment and completely zero sharding it to a lot of tasks with a very high reliability is still something that is... I I don't think there is a very strong precedent for that. And there is also, like, a trusted test of program for the Gemini robotics on device model where we work with a lot of different labs, and they have their specific robots. And I think the on device release showed, like, it controlling. This was one of my favorite parts of it. You can put very, very small number of examples. With 200 examples, you can learn a lot of different tasks on a lot of different bodies because you already have the foundation to do a very generic understanding of the physical world and also, like, how to control it from... And the data for how to control it with many different embodiments. I would think... I think a lot of the challenges with how the orchestration works is still split. There is... Surprisingly, there is the latency component. Right? Because when you have two models, like, one model is, like, doing the task, and another model is also, like, looking at how to switch the task. Like, when it is done and how... Now how do you progress to the next task? I think a lot of failure cases often also come from how to do the orchestration really well and... Yeah. So I think... And it is also compounding of errors. Right? Like, now if you have 10 tasks in sequence and now you have the ER model, which has a certain x success rate, and the VLA model, which has a y success rate. Now the... If you have a sequencing task, you are basically compounding the error. Does that make sense?
[52:55] Nathan Labenz: So if I ask a if I ask an AI robot to go fry an egg, how far does it typically get? Can... You know, are are we getting some successes on eggs fried, plated, and served to the table, or would that be beyond what one of your robots today could realistically do?
[53:18] Keerthana Gopalakrishnan: It depends on whether you have data or not. I think that if you there are ways to make, like, if you specifically just wanna fry an egg and how to do that reliably, there are algorithms that can make it very reliable just for that task. But if it's, a very generic model which has not seen how to fry eggs and now you bring it into your home, which is to be fair, I think models have made a lot of progress on visual generalization. They are now not distracted by, like, lighting conditions or the different setups of different homes and stuff. So I wouldn't... But I think it's like the task semantics and the context. Right? And the ability to make mistakes. I think especially especially for frying, it's like you can overcook an egg and no one would like that. And there, the room for error then becomes small.
[54:00] Nathan Labenz: Like Well, I overcook my own usually too, so that would... We could count that perhaps as a success for me.
[54:06] Keerthana Gopalakrishnan: Tasks that allow retries are much easier. Like, if you are doing, like, LEGO assembly, it's fine if you make a mistake. You can go back and fix it.
[54:14] Nathan Labenz: Yeah. Yeah. It's a great distinction.
[54:16] Keerthana Gopalakrishnan: If you crack an egg and then it's on the floor, then now you're in a mess.
[54:20] Nathan Labenz: Now you've got a mess. And you might slip in it too. What are the... So one of the value propositions of the Gemini Robotics two model, if I understand correctly, is it can work with any embodiment. What are, like, the most exotic embodiments that have come your way via the trusted tester program?
[54:40] Keerthana Gopalakrishnan: Pricelessly, not a lot of exotic embodiments. I think maybe the weirdest thing I've seen is people putting arms on drones and stuff. I think it... It's just that once you have... If you have data, you can learn anything. I think that is the principle. So I do believe that you can fit, like, very weird embodiments to it. But in general, I think the embodiments can look very like a normal distribution. Right? Arms are very... If you look at, like, the eigenvectors in the embodiment space, like, there are a lot of, like, bimanual arms. There are a lot of, like, hands. There are, like, humanoids. But... And then, like, everything weird is, like, a long tail in that sense.
[55:20] Nathan Labenz: Any kids' toys that you've seen that might be in demand this holiday season?
[55:26] Keerthana Gopalakrishnan: Kids' toys? Like, for robots?
[55:29] Nathan Labenz: Or Yeah. I don't know. I just can imagine. I'm surprised, honestly, that I haven't heard more demand from my kids just based on what their friends have and whatever for some kind of AI enabled little robotic toy that might, I don't know, honestly, walk around the house, but move around on wheels even and talk. And I'm not sure I want them to want this because I'm not sure I want it to be their life, but I'm surprised that I haven't seen more. So I'm just wondering if
[55:55] Keerthana Gopalakrishnan: the I'm also surprised
[55:57] Nathan Labenz: pipeline is there.
[55:58] Keerthana Gopalakrishnan: I think it's probably because the market is a little bit burnt. I remember when I was in school, there were a lot of companies trying to do my tabletop like a gaming robot. And I think... I remember when I was in CMU, there there was a lot of conversations around this. But maybe what a lot of people learned from that was like, people buy these things for novelty reasons, but they don't go back to using them. I don't know how that is. None of those robots were machine learned robots, which was connected to the large models, and now things are very different. I have... I I speak to Alexa and Siri and stuff, and sometimes I wonder if it would have been a lot more fun if they were much smaller.
[56:39] Nathan Labenz: I'm still waiting for good good Siri for god's sake.
[56:42] Keerthana Gopalakrishnan: Yeah. I've been talking to the Gemini, the live user interface when I've... I'm doing tasks or just talking about, like, my personal life problems and trying to get its perspective, like, I think you shared. And, yeah, it has been very useful in having an expert go over your problems where you're probably not very skilled to solve.
[57:03] Nathan Labenz: Yeah. I'm a big Voice Mode fan too, actually. In a previous conversation that we had, I asked you what you most needed from hardware makers, and you said, I need good hands. How are hands coming along? How is hardware coming along generally? Are you pleased with the progress? Is it gonna become a bottleneck for you relative to how fast the models are improving? Where do we stand?
[57:29] Keerthana Gopalakrishnan: When I last spoke to you, it was probably in the May... March of last year. Right? So here's the interesting thing. Right? We had... In the March of last year, we came up with Gemini Robotics one, which showed that a lot of dexterity with grippers. And in the summer of this year, and it's only, like, a year and four four months. Right? Three months. A year and a quarter. We showed Gemini Robotics two, which does multi finger dexterity, which can tie trash bags and stuff and do very fine control of multi finger hands. So it is... Here's where my internal model is also updating. Right? All the tasks that the gripper robots could do and that they were at the frontier of dexterity, the hand robots can do now. And the hand robots can do more things that the gripper robots could not do. So in a way, hands are now at the frontier of dexterity and grippers are maybe not. And I think this has been... I think I was quite surprised at how much the hardware has come along. There are a lot of hands on the market that are quite good that people are doing a lot of research with. There is still more work need to be done about how to make them very reliable, repeatable, and also make them cheaper.
[58:47] Nathan Labenz: How strong are they if my wife needs me to open a jar, would the robot be able to substitute for me and open the jar for her, or would it be like she would need to open jars that the robot can't open? Like, where does it fit on the kind of the, in addition to the dexterity, the ability to really torque something?
[59:08] Keerthana Gopalakrishnan: I think there is a spectrum, and it depends on which hands that you use. I think especially even off the shelf, there are a lot of variations. The Wuji hand is closer to, I think, maybe, like, a 10 year old, I would say. It's it's it's a bit weaker, and the sharper hand is... I think it can lift, like, 20 kgs. I don't know how much the torque stuff, but I've seen it open jars, so I'm pretty sure that's possible. The... Maybe not a very tight one. I don't know. So there is a lot of spectrum, and some hands are built for lifting more. The sharper hand is also much bigger than my hand, and the Wuji hand is maybe close... More closer in human sized.
[59:45] Nathan Labenz: It'll be funny if the last job for humans is opening stuck jars that the humans don't quite have the restraint for. Random aside, it's fine if the answer is nothing or it's all fake. But what is soft robotics? What do you know about soft robotics? What should I be paying attention to? I recently talked to a founder in this space, and, honestly, I came away more confused than I went into that conversation.
[1:00:10] Keerthana Gopalakrishnan: Wait.
[1:00:11] Nathan Labenz: The building. Basically, it was funny. They basically wouldn't really think. I was like, why did you schedule an interview if you don't wanna say anything at all about what you're doing? So it was a random kind of strange experiment or experience, but I did nevertheless come away thinking, what should I be paying attention to, if anything, in soft robotics?
[1:00:29] Keerthana Gopalakrishnan: Regards soft robotics, there is a wide spectrum, and I have seen things from people building soft bodies to people building very soft hands. And what's relevant for AI driven robotics is what's repeatable and what's durable. I think one trend that I'm very interested in is, like, the UMI type of effort for a lot of gloves and stuff and for a lot of, like, tactile modality. I think that type of how to do them really well and how to build that really well is still an area of research, and part of it is, like, soft robotics. Think, yeah, I'm not I'm not... I'm also not fully... Roboticists are of different kind. There are also, like, a lot of roboticists who are, like, very, like, very mechanical engineering focused and then who see the work that I'm doing. You're not a real roboticist. And I'm like, sure.
[1:01:19] Nathan Labenz: Okay.
[1:01:19] Keerthana Gopalakrishnan: Computer science roboticist and the mechanical engineering. Oh, you should go to a robotics conference. It's all kinds of people. Ikra is a good example. There's every kind of roboticist there.
[1:01:32] Nathan Labenz: I did... That sounds fun. I did have a chance to go to WAIC in Shanghai this summer, and there were a ton of robots there. I mostly felt like the demos seemed highly staged. I didn't feel like they were really demonstrating a lot of generalization ability in most of the things that they were showing off. It's like a CES scale show with multiple giant venues and every AI company in China basically there. It was a cool experience. The funniest robotic, specifically, experience I had there was sitting down at the Tencent booth. They had a sort of half humanoid. I don't think it had legs. I think it was just mounted upper body, and it was giving massages. So I got a very brief robot massage, some sort of like It pressure point was too short to really score. It didn't hurt me. It did apply some like nontrivial pressure, and I I walked away with a smile on my face. So I guess it wasn't too bad, but I'm not sure. I think I still would be highly confident that a human massage would be better. They're letting robots touch people over there now. It was... That's something. Speaking of robots touching people and potentially crossing other lines, Everybody, of course, this summer is talking about Open Face, the scandal, the incident, the... I sometimes call it a lab leak out of OpenAI. And it just has me thinking, like, maybe the research is all gonna work, and the alignment and the safety questions will ultimately be the bottleneck on robotics deployments. You could analyze that in multiple ways. One way to think about it is just like, when do you think if we put the alignment and safety questions aside, I might get my domestic servant robot, and then people can think for themselves. Oh, I think that's, like, before or after I would expect alignment to be good enough to actually want one.
[1:03:27] Keerthana Gopalakrishnan: Yeah. I think this is definitely an area that we think very carefully about. And I think of safety as, like, a capability. Right? Like, I... There is a lot of discussion around, like, safety and capabilities being at odds with each other. But people are not gonna use an unsafe robot and unsafe agents. I think the most useful agents are going to be ones that are safe that know what they need to do, and they're not overriding human instructions and causing chaos or breaking laws and stuff. So I think safety needs to be front and center in the design for AI and also robotics. And especially in robotics, like... In robotics, though, I think there are different types of safety. Maybe one thing that is not relevant to AI safety is operational safety. Right? Like, how do you make sure the robot is not, like, injuring people? And this is not injuring people because you're malicious, but just in... Not injuring people because you're dumb. And a humanoid being stable enough, like, there... It's not kill... Doing a wild plot to take over your house. But it's just if it falls, it can still be quite dangerous. So there are, like, a lot of different safety. So that's maybe, like, one type of safety resource that's, like, very specific to robotics, but not yet. I don't see it being a big thing in the AI space. But then there is also things like, what should I do if there are, like... Obviously, like, the higher level goal stuff. Right? Like, how do you make sure you do the task that the human is asking but without breaking guardrails? And then there is also the stuff about... And maybe this is also specific to robotics. Like, what happens, like, when some senses go wrong and how should I react in those situations? This one very funny demo that we showed was, like, the humanoid is doing something, and then someone goes in and puts a basket on top of its head. And now what you do, you can continue doing the task with your eyes covered, but the right thing to do there would be, like, really seeing that, hey. Like, something is obstructing my vision. Can I? Can you please remove it? So I think safety... And safety is like a full system thing. Right? Like, it goes all the way from, like, the design... The mechanical design of the robot system safety to all the way to, like, the ER and the high level brain. That's how I think about it. I do think as AI becomes more capable, like, we need to think... Spend more flops thinking about safety.
[1:05:49] Nathan Labenz: Yeah. Your example of somebody messing with the robot is an interesting one. And we've seen this with Waymo's and stuff, obviously, right, where people, like, go out of their way to cause trouble for AI systems.
[1:06:03] Keerthana Gopalakrishnan: In fact, it's something that I was very surprised of. It's like maybe because you... And it's it's hard. Right? You have two brains in you. One brain is the researcher brain who knows that they are machines, and now you you you don't really treat them as human like. But then there's the other brain, which is like, because they look so human like, they are... It's a bit too precidious. Right? Because when it's talking and stuff, you feel you feel kind of very human like around them. Now... So as a researcher, I'm, like, very careful a little bit. But I think when we see new people interacting with the robots, they very... People need to be of... More aware that these are machines. And I think sometimes when new people interact with the robots, because it looks very human like, it's very... It's that that boundary can blur. I also think this also creates a higher bar for humanoids. If you have a gripper based robots, like, it makes a mistake. Like, people don't expect it to be smart, but the humanoid just, like, grappling around and creating a lot of failures would be judged much harshly. Because even if you say, okay. This is the state of the manipulation research or whatever, people still expect humanoids to act smarter than robots that look anything unlike it.
[1:07:13] Nathan Labenz: Yeah. That's interesting. I don't wanna go down too much of a rabbit hole on consciousness, but I do think it's incredibly striking to me over the last year how much analogous structure has been found in LLMs. And personally, my willingness to entertain the possibility that AIs actually are somewhat conscious or sentient or something is up a lot just based on how much analogous structure we can now point to with things like functional emotions and functional well-being and the J space and so on and so forth. If you have any comments on that, I'm interested to hear them. If you wanna pass on that, you can just pass on AI consciousness for the moment.
[1:08:00] Keerthana Gopalakrishnan: Well, I I I wanna highlight one thing. So something that really happened for Gemini robotics two that was not in our previous releases is the human robot interaction component. So this is where it make like, very natural gestures, and it's not, like, preprogrammed. And previously, if you look at our videos from our last releases, it was... It would do all these nods and stuff where a bit more, like, preprogrammed HRI. And now it's, like, very natural HRI. And there, now you see, like, very funny things happening because the robot itself is deciding what gesture should I use as I talk to the person. And something that was very interesting was as we were, like, filming for the Gemini robotics too, there was... There are... If you look at the videos, there are also interviews with the robot. So people ask the robot, like, hey. How do you think you did the thing? And then the robot would be like, yeah. It was great, but maybe I made a mistake on this part of the task. And so it's... And and also we had the filming crew, and they were shooting the robot. And initially, the filming crew, they hadn't spent that much time around the robots. But at the end of it, they were like, action robot, and then the robot will do the thing. I think... I do think as we become more comfortable around humanoids and, and also as these models become very smart, I think, yeah, the lines are going to blur. I'll say one funny incident. So this was like the robot working in the the basement and the garage and then trying to pack things and stuff. And then there was an interview about it. What was the hardest part of you doing that task? And then the robot says, the videotape, which was my favorite object, I think it was very hard to put it in the in the basket instead of holding onto it. And then we're all sitting in the back, and we are all, like, like, laughing at the robot's response. And, yeah, it's kinda... Would a robot like a videotape?
[1:09:47] Nathan Labenz: I guess so. Self report. Self report is at least we... We're seeing somewhat reliable, although not always. Sure.
[1:09:54] Keerthana Gopalakrishnan: There have been some very, very adorable moments. Yeah.
[1:09:59] Nathan Labenz: Where are the gestures coming from? If they're not preprogrammed, what is it that is causing the model to be inclined to gesture to humans?
[1:10:10] Keerthana Gopalakrishnan: It is being prompted. There is a model that is using HRI to come up with gestures.
[1:10:15] Nathan Labenz: Gotcha. So, like, ER two is saying to the VLA model, raise your hand in a gesture as I give this response or something along those lines.
[1:10:25] Keerthana Gopalakrishnan: No. I think it's the... It's like following along in the conversation. Like, you can... It's not very explicitly prompted to do this or this, but it's like coming up with the gestures on the... So just like how the VLA is. Right? No one is asking the VLA, put your hand out and then go grab the thing. So it's more natural than that.
[1:10:43] Nathan Labenz: Fascinating. Would you call that emergent? Not quite emergent. Sounds like you're designed for it somewhat, but not super specifically. So it's semi emergent maybe.
[1:10:54] Keerthana Gopalakrishnan: Yeah. And this is also a field of research that we start thinking more about as you work on humans. So I think in the outside world, there's a lot of conversations around why use humanoids for... Why not just to arm robots and stuff? I think there are certain fields, entire fields of research that you would not touch because you would not be exposed to those problems until you look at what is the peak form factor that I wanna target. And I really like that maybe DeepMind is probably the only lab here in North America where you can study humanoid intelligence in a cross embodied way. It's not... We're not specifically building a brain for any one body, but it's very general. And there are multiple humanoids in our lab. So it's thinking about, like, how do you think about the intelligence substrate, but then also abstract away the body a little bit. And intelligence is, in some ways, a little bit abstracted away from the final form factor. And there are multiple fields that only arise that when you do different humanoid bodies. One is human robot interaction becomes a much bigger thing. Full body control becomes a big... Bigger thing. Multi finger dexterity becomes a a a new front. And to get exposed to all these three frontiers and research problems, you need to work on it.
[1:12:12] Nathan Labenz: So one big question I have about kind of the future of the field... Shout out to doctor Jim Fan from NVIDIA for an outstanding talk he gave that inspired a couple big questions for me. He makes the case that, basically, model... The robotics models in particular need to think ahead. Right? They need some sort of forward looking simulation, which is not exactly synonymous with world modeling, but certainly highly related to world modeling. But the key idea I think that he has, going back to my earlier question on the response time of a classic reasoning LLM backbone system is you can't... In his telling, you can't just have the model look at the present all the time, reason about the present, take action, and run that cycle over and over again. You need to have some way of kind of projecting into the future how things are gonna go so that I can move and measure the delta between my expectation and what's happening. And that's like the path. Certainly, that seems to be from what I understand of how our brains work. That seems to be what we're doing. Do you think... I think this is... You could look at LLMs and say, they've come awfully damn far on a token by token prediction without explicit world models. Maybe robotics can do the same. Maybe it's different. Where do you come down on that? I'm sure pretty central debate.
[1:13:35] Keerthana Gopalakrishnan: Yeah. I think the jury is still out there. Right? Ultimately, the thing about working as a roboticist is that you are not emotionally attached to one method or the other method. You're emotionally attached to the problem And whatever way you solve the problem, you will pick that. And, yeah, there are many schools of thoughts. Definitely, world modeling is a very interesting area for robotics and not just for as policies, but also simulation and other things. And I think that you... I I say the jury is out there still because there are obviously the the VLAs, the VAMs, but also look at how the agent robotics are going and think we... We'll see. And this is maybe one of the most exciting parts of working in robotics is that it is still very early that the recipes haven't stabilized. Different people can have very different schools of thoughts about how I'm gonna go about and solve robotics. And we'll see who's right, and all of them could be right. And it might be the future until we have solved the problem. It's very hard to very definitively say this is the solution or this other thing is the solution.
[1:14:39] Nathan Labenz: Yeah. Interesting.
[1:14:40] Keerthana Gopalakrishnan: Okay.
[1:14:41] Nathan Labenz: So it sounds like you're open minded, and, obviously, Google DeepMind has world modeling efforts that are pretty well established as well. So to the degree that you need to pull a world model off the shelf and to start simulating into the future, you have a pretty good foundation right at hand to do that kind of work. He had a couple other interesting points I'd be interested in your take on. One was basically he thinks that we are gonna go to egocentric video as kind of the main data used for training. He thinks that, like, teleop will go to a a tiny scale or tiny portion of data and even the sort of hand. We've seen a lot of innovation. I think it's been pretty cool around just like wearable robot hands that you can use, but he thinks even that will ultimately just give way to first person POV videos. And then beyond that, it's just gonna be simulation, and he thinks basically like all the problems of simulation will just get solved, and that will really be the way that we'll get to enough robotics data to solve robotics. Any thoughts on how that data... That vision for the future of data compares to yours?
[1:15:56] Keerthana Gopalakrishnan: I think everything needs to be backed by results. Right? And it's it's... I don't think there is much point in making a prediction. And, usually, you have a prediction, you run an experiment to validate that prediction, and then now you know that is the right prediction. And it... These things can also change over time. And I think as Jim suggested, it's great that he has very strong views on certain types of data. I think it might also very much look like it's a mixture over time. Like, you look at LLMs, they're trained on all kinds of data, and different data has different purposes. Like, the teleop data is definitely useful in grounding into the robot's controls. Umi data is definitely very useful in learning much scaling scaling without robot in the loop. So much cheaper scaling. And human data is also very... Like, egocentric human data is also very useful in getting broad semantic understanding about how... What to do and stuff, but they also have their failure modes. Right? If you imagine that you controlling very fine hands, you don't get very fine signals about where the hands are unless you're getting... You're estimating all that from vision, and vision has a lot of errors. And so, like, each data modality has different strengths and different drawbacks. Like, the UMI data, like, with sense... With sensors, you are getting the action labels, not from vision, but from sensors. So it's a bit more accurate. So there is... I I would imagine each data source is on on the excesses you can plot, like, scale and the precision. And maybe, like, teleop, not very scalable, but highly precise because it's exactly the robot's data. But it is also not very, like, future proof. Right? Because as your robot evolves, your teleop data starts becoming little less useful even if you have very good cross embodiment. Then there is UMI data, which is, like, more scalable, but because of all these sensors, it's more precise than human data, but it's less scalable because of all these sensors. And then there's human data, which is, like, very scalable, but can also be very noisy because you don't have very precise and effective actions. And you also need to, like, now put that on the robot. So there are errors that come from... And each human is very different. Some humans are small, some are large. And due to that, how they do the end effective tasks is very different. So I think it's going to be, like, probably closer to a mixture. And the thing is hardware is changing each of these paradigms differently. Right? Like, as very good only hardware comes on the market, like, that... How to scale that now goes up. So I don't think it's wise to take a very principled view about this because it... All of this will change as new embeddings emerges and better hardware emerges. And you probably need all of all types of data.
[1:18:48] Nathan Labenz: Then maybe one other aspect of his talk that I thought was interesting. It wasn't a main focus, but it got me thinking a little bit, was he used the term physical API. And it got me thinking, like, how do I want to interact with robots in a world where I'm interacting with, like, multiple different robots? Do I want them each to present as a entity unto themselves? Or maybe it turns out to be more of like a her model where I have one entity that I, like, interface with the most, and it kinda deals with all of the robots as a fleet. And therefore, the robots are not like individual entities in my life, but they're all sort of extensions of this, like, single personal superintelligence or whatever you wanna call it that can run all these other things. Do you have an intuition for how you would want to interact with many robots in your life?
[1:19:46] Keerthana Gopalakrishnan: Yeah. I I don't have the context of this physical API conversation, but it also ties a little bit into what we were discussing before. Right? I think the surface area of our interaction with robots needs to be highly multimodal. And I, the user, needs to have both high level and low level access to robots. I think that is probably the only way that we can build towards a very safe and flexible future. And, anyway, from language and if robots code and if a task requires multiple robots, then I need to be able to hand it off, and then it can handle the collaboration. I think the Gemini robotics too kind of had this where both the dual robot and the Apollo was doing the task. One was doing one and then it stops and the other one takes over. So there is... I think... Yeah. I maybe... I don't know. This is a boring answer which I've said before. It's probably going to be a spectrum. Right? If there are tasks which require multi robot collaboration and a lot of agentic swarms, and then there, I need to be able to hand off and they need to figure out how to do things. And then there are tasks where I'm only interacting with one robot, and I either want to give it text instruction because I'm just talking to it, or I wanna give it more fine grained specific feedback. And there, I need to go even more global level. I think for some, a lot of safety and interpretability reasons, maybe we need to still maintain much lower level access to robots.
[1:21:13] Nathan Labenz: Yeah. What you're saying there reminds me of what I'm increasingly thinking of as, like, the universal UI for software products, which is, like, there's the product itself, and I can click on it and edit fields and drag stuff around. And then there's an agent or assistant that kinda sits alongside that and has basically all the same affordances that I do. And most of the time, I probably just wanna tell that thing to do it, and I don't really wanna concern myself with the details. But I also have found that it's really annoying sometimes if the app doesn't let you get to that level of detail because sometimes the AI is just, like, not getting it on its own. So I can see a way that paradigm ports over to robotics.
[1:21:56] Keerthana Gopalakrishnan: Yeah. Especially, let's say, like, I cannot tell a robot to clean my desk because I like my desk done in a certain specific way. So I need... I would like the ability to give it more feedback and even personalize it to my own, like
[1:22:14] Nathan Labenz: What's up with these days with... For lack of a better term, I'll call it circuit breakers on robots. What I mean by that is I saw one example where the company had made a sort of inflated airbag in the torso of the robot. The idea was like if you poke that thing, it is immediately disabled. The pressure of the bag that you could easily puncture as a human immediately disables it. I thought that was just interesting new idea in terms of, like, how you could, worst case scenario, have a... An an off button locally. But more broadly, there's, like, what happens when it encounters resistance it's not expecting or whatever. How has that evolved over the last year plus?
[1:23:02] Keerthana Gopalakrishnan: Yeah. So I think for the longest time, this air bubble that you break, there has been a version of that. It's called e stop or kill switch. So robots... Most robots have a button that you can press, and then it will depower. And there are many phases of it. Right? There is the hard e stop where you stop all the kind of power to the body, and then it can just collapse, which can then be dangerous on its own. There is also the soft east of where you just freeze the robot, and this is often safe for bipedal and other platforms, which are not stable by themselves. And in recently, I think, especially in manipulation, there are a lot of new features kind of emerging to control forces at the end effector. Right? Imagine that you wanna stack two chips together. You can... If you press too hard, you crumble the thing. And... Or if you are handling very delicate material, you need force control and backdrivability at the end, compliance there. So I think sensing is improving quite a bit that you can have very... You can read what the end effector forces are, and then the model can take much better decisions than they would if they didn't have that information.
[1:24:10] Nathan Labenz: Last question. Kind of a galaxy brain or maybe a fly brained one. So there's been all these demos, I'm sure you've seen recently of people using the fruit fly brain to do all kinds of things, including seemingly to control some small robots.
[1:24:26] Keerthana Gopalakrishnan: Actually, not seen them. Maybe our, like, Twitter feeds are still fairly personalized.
[1:24:32] Nathan Labenz: There's a... Yeah. There's one... I'll send it to you. There's one where somebody has trained this thing to control a little, I think it was called a Beastrand, one of those, like, beach walker robots, but a real small one. The connectome of the fruit fly brain was fully mapped and put online in digital form, and people have just started downloading it and training it to do different things. I'm not a 100% sure how real it all is, but there's enough of it going on that I think some of it's real. It sounds like you don't have necessarily take on that. I'd welcome any comments, but maybe zooming out, I wonder far future, do you envision yourself with some sort of robot body? Is there like an... A neural brain computer interface where you imagine having a hybrid form yourself? Yes.
[1:25:21] Keerthana Gopalakrishnan: Yes. And it already exists in some sense. Right? Like, when people are teleoperating, their physical body is somewhere, but their robot body is the one... Just teleop itself is like that. And then there is people who are building remote teleop that works across The Atlantic and stuff. So the person is sitting in Europe, but the robot is in in The United States, like, stacking, I don't know, stocking shelves or something. So definitely, this is very interesting. I would imagine someday that you get to explore Mars or something with robots because you are teleoperating them. Yeah. I think the future is exciting. One of my friends, the thing that he wants is he wants to walk around. He wants, you you know, the guards in India, they have, like, all these arms popping out from their behind. But now you can build them. Right? You you can put those trust and arms on your back, and then those arms can go pick things. And you can just walk around and assess with those. So that's the thing that he wants to build as a hobby. I also think about, like, my dog. I looked up how many parameters she has in our neural network. Apparently, it's 25,000,000,000,000. So she has more parameters than lot of the models out there even though her her language skills are not that developed. Yeah. I think the future is very interesting.
[1:26:35] Nathan Labenz: Yeah. That is a modest way to put it, I'd say, and that might be a great note for us to end on. As always, this has been a fantastic conversation, and I appreciate your help in catching up with everything that's going on in robotics. Kirtan Agoplakrishnan, thank you for being part of the Cognitive Revolution.
[1:26:53] Keerthana Gopalakrishnan: Thank you so much for having me.
Outro
[1:30:05] If you're finding value in the show, we'd appreciate it if you'd take a moment to share it with friends, post online, write a review on Apple Podcasts or Spotify, or just leave us a comment on YouTube. Of course, we always welcome your feedback, guest and topic suggestions, and sponsorship inquiries either via our website, cognitiverevolution.ai or by DMing me on your favorite social network. The Cognitive Revolution is part of the Turpentine Network, a network of podcasts which is now part of a sixteen z where experts talk technology, business, economics, geopolitics, culture, and more. We're produced by AI Podcasting. If you're looking for podcast production help for everything from the moment you stop recording to the moment your audience starts listening, check them out and see my endorsement at aipodcast.ing. And thank you to everyone who listens for being part of the cognitive revolution.