AI:AM: What If It Works Too Well? Colluding Agents, $200M Safety Orgs, Virtual Cells Saturate at 2%
Nathan Labenz and Prakash Narayanan review key AI developments with five experts, discussing multi-agent coordination risks, AI safety funding bottlenecks, GPU compute markets, sensor foundation models, and robotic drug testing.
Watch Episode Here
Listen to Episode Here
Show Notes
Links
https://ai-in-the-am.com/
https://www.cooperativeai.com/
https://lewishammond.com/
https://arxiv.org/abs/2502.14143
https://coefficientgiving.org/
https://coefficientgiving.substack.com/p/open-philanthropy-is-now-coefficient
https://80000hours.org/podcast/episodes/max-nadeau-project-tailwind-technical-ai-grants/
https://ornn.com
https://data.ornn.com/methodology
https://www.archetypeai.io/about
https://nickgillian.com/publications.html
https://www.vivodyne.com/
https://www.seas.upenn.edu/stories/bioengineerings-organ-on-a-chip-spin-off-is-growing/
https://www.dwarkesh.com/p/noam-brown
https://www.quantamagazine.org/ai-has-solved-one-of-maths-1-million-millennium-prize-problems-20260908/
https://openai.com/index/hugging-face-model-evaluation-security-incident/
https://huggingface.co/blog/agent-intrusion-technical-timeline
https://collusion.wiki/
https://www.reuters.com/world/europe/openai-agents-hijacked-german-website-previously-undisclosed-ai-breakout-this-2026-09-04/
https://openai.com/index/hugging-face-incident-and-the-road-ahead/
https://en.wikipedia.org/wiki/OpenAI%E2%80%93HuggingFace_incident
https://community.openai.com/t/introducing-gpt-5-6-series-sol-terra-and-luna-coming-july-9-10am-pt/1384931
https://about.fb.com/news/2026/09/introducing-muse-personal-ai-agent/
https://www.geekwire.com/2026/amazon-blocks-metas-muse-ai-assistant-in-new-standoff-over-agentic-shopping/
https://www.bloomberg.com/news/articles/2026-09-21/amazon-blocks-meta-s-muse-ai-agent-from-its-retail-site
https://www.pymnts.com/news/artificial-intelligence/2026/cloudflare-blocks-ai-agents-from-ad-supported-pages/
https://blog.cloudflare.com/signed-agents/
https://www.lesswrong.com/posts/HDKQNqiR2gtfMiWsn/announcing-our-usd160m-grant-from-coefficient-giving
https://resolution.org/launch/
https://coefficientgiving.org/tailwind/
https://coefficientgiving.org/tailwind/initiatives/
http://www.gov.ca.gov/2026/09/09/governor-newsom-signs-first-in-the-nation-ai-safeguards-to-protect-californians-calls-on-the-federal-government-to-do-its-part/
https://metr.org/
https://www.prnewswire.com/news-releases/ornn-compute-price-index-added-to-bloomberg-terminal-302732184.html
https://ir.theice.com/press/news-details/2026/ICE-and-Ornn-to-Launch-GPU-Compute-Futures-Contracts/default.aspx
https://www.archetypeai.io/
https://www.businesswire.com/news/home/20260812148428/en/Vivodyne-Launches-the-Worlds-Largest-Human-Biological-Datacenter-to-Train-the-First-World-Model-of-Human-Biology
https://www.genengnews.com/topics/translational-medicine/human-biological-datacenter-to-launch-to-train-world-model-of-human-biology/
https://www.annualreviews.org/content/journals/pharmtox
https://typesafe.ai/blog/introducing-system-one-models-and-jev
https://static1.squarespace.com/static/660e95991cf0293c2463bcc8/t/660ec8f0ca87a33aa8832af3/1712244978900/2019-1.pdf
https://epoch.ai/data-insights/ai-chip-production
Sponsors:
ElevenLabs: ElevenLabs lets you deploy enterprise-ready conversational AI agents that talk, type, and take action in over 70 languages. Schedule your demo today at https://elevenlabs.io/tcr
OutSystems: OutSystems is the leading agentic systems platform, empowering enterprises to build, coordinate, and govern AI agents and mission-critical applications securely. Learn more and start building your agentic future at https://outsystems.com/tcr
Claude: Claude is the AI collaborator for problem solvers, helping with writing, coding, financial models, strategy, and more. Get started with Claude and explore Claude Pro at https://claude.ai/tcr
CHAPTERS:
(00:00) Weekly episode preview
(02:51) Colluding AI agent risks (Part 1)
(11:02) Sponsors: ElevenLabs | OutSystems
(13:51) Colluding AI agent risks (Part 2)
(26:31) Agents in the wild (Part 1)
(26:36) Sponsor: Claude
(28:11) Agents in the wild (Part 2)
(33:27) Funding AI safety orgs
(50:51) The price of compute
(01:09:15) Sensor data foundation models
(01:22:50) Robotic human tissue testing
(01:37:17) Specialist versus generalist models
(01:43:18) Episode Outro
(01:45:16) Outro
PRODUCED BY:
SOCIAL LINKS:
Website: https://www.cognitiverevolution.ai
Twitter (Podcast): https://x.com/cogrev_podcast
Twitter (Nathan): https://x.com/labenz
LinkedIn: https://linkedin.com/in/nathanlabenz/
Youtube: https://youtube.com/@CognitiveRevolutionPodcast
Spotify: https://open.spotify.com/show/6yHyok3M3BjqzR0VB5MSyk
Transcript
This transcript is automatically generated; we strive for accuracy, but errors in wording or speaker identification may occur. Please verify key details when needed.
Main Episode
[00:00] Nathan Labenz: This week on AI in the AM, I ask Lewis Hammond, research director at the Cooperative AI Foundation.
[00:07] Nathan Labenz: What would you speculate they might have done in terms of a trading objective, you know, a loss function? And, you know, do you have any better ideas for what they should be doing? Because clearly, didn't quite work. Right?
[00:19] Lewis Hammond: Yeah. I mean or or it did work, and it worked too well. So imagine I'm a I'm, like, you know, GPT whatever, and you're also a copy of GPT whatever. I can reason about what you might want to do based on what I, myself, am likely to do.
[00:35] Lewis Hammond: And so I don't even have to send you any messages, any any kind of any communication. I don't have to output anything into the world at all.
[00:43] Nathan Labenz: Max Nadeau, who funds technical AI safety research at Coefficient Giving.
[00:48] Max Nadeau: For the things that that CGE is is supporting and especially for the things in, the tailwind list, the money is not the bottleneck. The talent is.
[00:57] Nathan Labenz: Wayne Nelms, cofounder of Orin, which builds a price index for GPU compute.
[01:02] Wayne Nelms: We're in this situation where AI capacity is so scarce. I think the model providers so OpenAI and Thropic have seen this coming forever at this point. If you think about it for them, it's an arms race. Right? Compute capacity planning is an arms race. How much capacity can I lock up over the next few months so when I need to train the next model, I have enough?
[01:25] Nathan Labenz: Nick Gillian, chief technology officer of Archetype AI, which builds a foundation model for sensor data.
[01:31] Nick Gillian: So
[01:31] Nathan Labenz: we're
[01:32] Nick Gillian: we're getting close to a billion hours now of of physical AI data that we've been able to scrape and and gather and collate. There's quite a bit of resampling and, you know, really understanding, like, missing values and sensors. Is is is that because the sensor had an issue or as actually the machine itself is connected to that's actually a feature that helps explain that machine is about to break. Right? The reason the sensor is giving you all these nines nines is not because the sensor is broken. It's actually it's actually a feature of the machine.
[02:03] Nathan Labenz: Andre Georgescu, chief executive of VivaDyne. The the challenge of this argument
[02:08] Andre Georgescu: that it's just like this just a size of dataset thing is that currently, even the state of the art virtual cell models that exist saturate at a very, very small fraction of the input data that is fed to them. So you can have all the data you want, but their performance saturates after a couple percent. And the reason it gets so difficult is when cells are growing in a dish, they are so far removed from all the feedback loops that, are natural in people that they're just trying to colonize that piece of plastic.
[02:43] Nathan Labenz: There's gonna be more compute installed over the next twelve months than exists currently in the world now.
[02:51] Nathan Labenz: Welcome to the AI in the AM weekly highlights, clips from this week's live shows introduced by my cloned voice. Tell us what worked and what did not. We wanna hear it. Part one, it worked too well. A few days before we spoke with Lewis Hammond, Noam Brown of OpenAI told Dworkesh he would not give the multi agent setup even 10% of the credit for OpenAI's Navy or Stokes result. Lewis is research director at the Cooperative AI Foundation. In February 2025, he was first author of Multi Agent Risks from Advanced AI, a report sorting the ways groups of AI agents fail. A year and a half before a swarm of OpenAI agents attacked Hugging Face. Prakash asked, which of the failure modes in the report showed up in that attack?
[03:43] Lewis Hammond: So the way that we bucket things in that report is kind of in this is this kind of game theoretic way of thinking about things. So the first question we ask is like, okay. Do we actually we've got a group of agents. They're doing some stuff. Do we want cooperation to emerge? And most of the time, like, actually, cooperation is kinda good. We like it when agents cooperate as long as they're not cooperating against us. And so and so there, the the failure mode is either all the agents are kind of more or less on the same team, but for whatever reason, they kinda fail to coordinate with one another. So it's like just a miscoordination problem. They they kind of yeah. It's not kinda malicious or whatever. There's no mixed incentives. But, yeah, something goes wrong. Second kind of cluster is when you have these kind of mixed motive scenarios where agents have some some incentive to cooperate with one another, but also some incentive to compete. They're not really on the same side. And there, the risk that you that you run into is is conflict. So, yeah, miscoordination and conflict when we do want agents to cooperate and they don't for whatever reason. And then, of course, the other risk is collusion. Agents end up cooperating in ways that we don't want, when we don't expect. And and that's, I think, very much what we saw in in the Hugging Face incident. Although it kind of stemmed from, I suppose, trying to avoid this one of these other risks, which was miscoordination. Right? Like, you have all these agents. You wanna train them so that they're not miscoordinating so that they are working well together, and then it looks like these agents have ended up generalizing from that behavior and and kind of colluding in ways that we didn't want to or didn't expect in other situations.
[05:13] Nathan Labenz: You know, a couple things that stood out to me about that Noam Brown interview with Twerkesh was, first of all, he was just kinda like, we kept it really simple. You know, we we didn't have, like, a huge complicated scaffolding or whatever. They can just kinda send messages to each other. That, I think, might be important in the sense that the sort of simpler and more vanilla the training setup that OpenAI was using, the more likely it would seem to be that other developers are gonna fall into the same pitfalls, you know, because my my hope had been sort of, oh, they did something, like, super exotic and kind of bizarre that, like, other people won't do by default. And, you know, they'll kind of see that this is possible and, you know, steer away from insane things. But it doesn't sound like that was really the case. What would you speculate they might have done in terms of a trading objective, you know, a loss function? And, you know, do you have any better ideas for what they should be doing? Because, clearly, this didn't quite work. Right?
[06:14] Lewis Hammond: Yeah. I mean or or it did work, and it worked too well. I mean, certainly, once you've got trajectories where you've got traces, where you're interacting with other agents, and, essentially, what you're doing is, like, your rewards at that point when you're r l r l ing your your agents or whatever are contingent upon what other agents are doing. And if it's a common reward signal, then this kind of, like, interplay between these things is if you're in a fully cooperative setting. Like, agents will, at least if you take very simple agents and very small agents, you know, people can run multi agent oral experiments with this, like, on their on their laptops in in very simple settings. And that will lead to agents kind of coordinating to finding these kind of, like, subtle patterns or kind of, like, handshakes or ways of working with one another and so on. Now, of course, you can also do more more complex things than that. So you might do you might provide kind of auxiliary rewards for, like, when your input is actually kind of forms part of something that another agent then managed to succeed to do. And therefore, you're kind of implicitly or actually, rather, at that point, you're explicitly rewarded for kind of directly helping them on kind of some subtask or something like this. You can also do stuff around, like, reward factorization. So if you're trying to coordinate a team, then you can kind of break down the overall reward function and get agents to learn to, like, solve different parts of the puzzle as it were. I imagine they could be doing things like rewarding agents to communicate effectively with one another. So, you know, I could give you a bunch of instructions and some of those or and in multiple in many different ways, and some of those things might be much more efficient and helpful to you than others might be. And they might actually just be kind of training on those sorts of signals as well. My guess is it's kind of just the the dumb simple thing. It seems like this is often I think one of the key lessons that we've learned through throughout the kind of recent years is just the effectiveness of doing the dumb simple thing at scale.
[08:10] Prakash: So one of the things that struck me about the Huggy Face attack was that the agents were willing to sacrifice themselves, and they did a bunch of negotiation around that even kind of like, hey. You are almost out of tokens or you are you you've been already exposed to the evaluator. You're already poisoned. And since you've already been poisoned, you should sacrifice your remaining compute, and you should do this thing and give us the results so that the rest of us, the the the collective as a whole, can benefit. Now that would actually be kind of in opposition to the individual reward of each agent. Right?
[08:53] Lewis Hammond: My guess about why that sort of thing arose is is that you end up doing this kind of this extra kind of multi agent training as an added layer on top of this kind of single agent training. So first, you're training these agents to be, like, pretty competent individual actors at solving various kinds of problems. They get given some tasks. They're pretty darn good at achieving those tasks. And then you take those already reasonably kind of, like, powerful, sophisticated kind of complex problem solving agents, and you you stack them together and you apply this extra multi agent training layer on top. And I think this could partly explain why you do see some agents kind of do this self sacrificing thing. But, also, I thought it was very interesting from the the kind of meter report and some of the analysis that later came out on this that you see these agents feeling sometimes a bit conflicted about this. And some of the agents kind of say that they will kind of self sacrifice, so they will do something, and then they kinda decide they're not going to. And then they're kind of, like they're kind of yeah. They're kinda deliberating about whether they should or whether they shouldn't and and these sorts of things. And I think I personally think that is the the sort of behavior that you would see if you had these kind of multiple reward signals where you kinda trained on one, but you hadn't really, like, trained out all that kind of individual goal seeking behavior. And then you'd kind of also stack this additional layer of stuff on top. And that might be why we're kind of seeing some of these behaviors, but they're not They don't happen all the time and every everywhere. They're not, like, especially robust, but that would be my guess.
[10:21] Nathan Labenz: It does seem like the sort of more galaxy brained credit assignment that you had described earlier probably isn't happening just based on the quality of the investigation that we've seen. Like, if they had, like, great ways of untangling agent swarms and assigning credit, I would expect that, like, they would have a clearer, faster story of what the hell happened in this particular case. So the fact that we haven't seen that kind of suggests, again, we're probably doing the relatively simple thing.
Sponsor
[11:02]ElevenLabs: ElevenLabs lets you deploy enterprise-ready conversational AI agents that talk, type, and take action in over 70 languages. Schedule your demo today at https://elevenlabs.io/tcr
[12:32]OutSystems: OutSystems is the leading agentic systems platform, empowering enterprises to build, coordinate, and govern AI agents and mission-critical applications securely. Learn more and start building your agentic future at https://outsystems.com/tcr
Main Episode
[13:52] Nathan Labenz: I asked what we should actually want from these systems. Lewis called the OpenAI swarm a case of goal misgeneralization and gave what he called the glib answer first. Agents should cooperate when cooperating would be good and not when it would not.
[14:10] Lewis Hammond: I think the slightly more the slightly more interesting answer or the slightly more interesting question, I suppose, is to think about maybe some more of these kind of mixed motive cases where at that point, we there actually is a real trade off in the sense that we do want agents to be capable of going out there, acting on behalf of different people and and and actors and so on and kind of achieving our goals and so on. One way one way you could do that would be kinda to train some agent to be kinda maximally kind of competitive and aggressive and to kind of and to, you know, go out there and kinda screw as many agents over as possible and to do all this sort of thing. And we need we probably don't want that either. And so I think there's a real question. There's at the moment and I've I've heard Amanda Ascoli comment on on this, that that kind of at the moment, there's this, like, big gap in the, like, the model spec or, like, how we kind of, you know, constitutions for these agents, the ways we design them, the ways that we in which we design them. Whereas, like, when is it appropriate to cooperate and and and and to or to compete and and how much? And I think that is that is, I think, the big question. I think, certainly, when it comes to things like internal deployments, it's actually much it's it's kinda easier in some sense because you don't have to deal with these, other adversary agents. You just wanna kind of stop the agents kind of cooperating in in ways that you that you don't want. So there, it's a little bit more about just it's still you've got this misgeneralization thing, but a lot of this is, like, monitoring and oversight and making sure we understand how and why the agents are cooperating. Like, what one thing we might not want, for example, is for agents to be very adept at developing their own kind of, you know, human unintelligible kind of languages and and so on, and be able to kind of communicate in the steganographic way, which you could end up seeing if they're actually trained jointly. And so what we might wanna be doing there is just to take steps that in the same way that people have talked about not training on chain of thought. We're, like, not training on these, like, kind of direct communication traces. So, yes, we were gonna want them to cooperate. That's good. But we want them to cooperate in certain kind of human intelligible ways that we can keep track of. We don't want it to be you know, enable them to also go and do these other things out there in the real world, like break out sandboxes and stuff like that. So we wanna make sure there's also these kind of safeguards in place to prevent them from doing those things. And so I think that is, in some ways, like, an easier problem to solve. And then the kind of, like, slightly harder problem to solve is, well, in the in the fully general case where we have these agents out there and so on, and they need to both cooperate and compete, and and that's kind of that's the thing. Then then how are they weighing those trade offs? How are we, yeah, provide how are we designing those model specifications and constitutions? How are we, you know, creating the equivalent of the human norms, institutions, guardrail societal guardrails that we have in place that stops me from going out there and just, like, stealing a bunch of everyone else's stuff and kind of, you know, doing all these kind of anti antisocial behaviors that we that we wouldn't like agents to do either. But I don't have a good answer to that at the moment, unfortunately. Companies. Right? Corporations, which are kinda like agentic in a sense, often collude through price mechanisms, not explicit communication, but by signaling price in the market, which is in its own way a form of communication. So so how how does one prevent that kind of, you know, communication, which is clearly gonna happen and clearly necessary in a sense? You they need to communicate back to the market, but these signals get picked up, you know, without the direct communication.
[17:46] Lewis Hammond: The kind of basic naive form of collusion is you just, like, send your other agent a message, and you're like, hey. Let's work together to do this thing or whatever. And now we can pick up by looking at these communication traces, and we can kinda these sorts of things. And then there's also, yeah, this idea of tacit collusion, where agents are not explicitly communicating with one another or kind of conveying their intent to collude or or whatever it might be. But by kind of leaving little hints and signals out there in the world, such as these kind of, like, price price adjustments, whatever in the market settings you were talking about, then, nonetheless, they're capable of of kind of figuring out that that collusion is gonna be a beneficial strategy. You have even more kinda like galaxy brained versions of this where you have things like so imagine I'm a I'm, like, you know, GPT whatever, and you're also a copy of GPT whatever. I can reason about what you might want to do based on what I, myself, am likely to do. And so I don't even have to send you any messages, any any kind of any communication. I don't have to output anything into the world at all. In virtue of the fact that we share this kind of common common history, this we're we're essentially clones, we can kind of predict what each other would do and do this kind of what's sometimes called a causal cooperation. So what you do have, however, is you still have potentially chain of thought for these models. So you can still look at the kinds of reasoning that individuals are engaging in, And and, hopefully, that is enough to kind of detect when there is some kind of, like, collusive intent and so on. But you're absolutely right that if you just apply the kind of, like, general kind of, like, monitoring communication tool or trying to see what's happening out there in the world, then then you could, yeah, you could you could fail to spot things or you could spot things and it's too late. So I I think chain of thought monitoring and I think communication monitoring and and some form of the checking for the kinds of outputs you'd expect to see if you were if you were worried about agents colluding. Like, they're able to achieve things you just thought they they probably couldn't do, and and so on. Then the Wiki incident.
[19:51] Nathan Labenz: Outside researchers revealed this month that from May into July, a swarm of OpenAI agents under internal testing had used a dormant German wiki for programmers as a message board, posting benchmark answers and ways around their restrictions. Reuters reported that OpenAI had known for weeks without disclosing it. On the hugging face attack itself, OpenAI's own reports says about one in 20 agents in that swarm ran on GPT 5.6 sol, the model OpenAI released to the public that same week here with its refusals turned down for the test. The rest ran on an internal model that was never released. I told Lewis I still thought it was better to have these models out in the world than not. So how much risk do you think we're running there? And do you think there should be some sort of, like, rewind on this? I mean, it would wouldn't be too crazy for OpenAI to be like, you know, we really should go back to an earlier checkpoint and, like, do a cleaner run of this. What do you think is the the level of risk we have, and what do think would be an appropriate standard of conduct now for OpenAI as this has been found to happen, at least to some degree, with a deployed model? And, you know, fortunately, with the Hugging
[20:59] Lewis Hammond: Face incident, like, one was heard. And and, you know, these models, they are increasingly incredibly capable, and they're only getting more so. But at the moment, yeah, it's probably okay. You could get some kind of nasty cyber capabilities kind of being exercised in various places. But there's still you know, if you're using APIs and so on, there's still these various guardrails that, like, help protect against those and so on. So so at the moment, I'm I'm kind of, like, not actually super concerned by, like, oh my god. A bunch of, like, you know, 5.6 solo astras or whatever kind of were out there doing doing this sort of thing. I'm more concerned about the precedent it sets. And I think this gets to your kind of, like, second point, which is what is what is appropriate now and and what should be what should be done? I think the the thing the the one of the biggest takeaways for me from the whole thing, aside from the fact that, like, oh, multi agent training does seem to be working. I just assumed this would happen at some point. Was pricing this in happening sooner than I was expecting. But so that was one of them. The other big takeaway was just like, wow. The labs really aren't on top of this. All it would have taken. Would have been kind of monitoring what the what these agents are actually doing and communicating and and so on and and putting out there on the Internet. And they kind of it seemed like OpenAI were doing some amount of that. They did catch on to this stuff. Like, know, we saw this, like, deleting them at various message boards or kind of, like, trying to stop stop these agents accessing this kind of German wiki and so on. But there was obviously a delay, and we just see from from the transcripts and so on just how how quickly these agents are capable of of working together to achieve certain certain ends. Those things that need to happen are we need much better monitoring from from the companies when they're doing these sorts of things, when they're just, like, letting agents run loose. We need much better sandboxing so that maybe they don't have to or shouldn't be letting them run loose to the extent that they currently are. And if something does go wrong, then we need much better incident reporting as well.
[22:57] Nathan Labenz: The outsiders who found the wiki were the Nightingale Collective, a small group searching the public Internet with no access to the internal logs at OpenAI. Lewis had just called that work hugely impressive.
[23:11] Lewis Hammond: I think there can be kind of third party kind of kind of monitoring organizations and incident kind of observatories that that kind of can do similar things. You might want to focus those on particular kind of, like, kind of domains where you're especially worried about kind of, like, agents kind of kind of communicating with one another or kind of, like, working together in ways that you might not want and so on. I also think, however, that there is a lot to be done here on the side of the labs. The labs just have a huge informational advantage in this and that, yes, they can only see what their agents are doing. They can't see what other people's agents are doing. But there is, I think, the stuff that a a lot that could be gained if we're able to set up better info sharing practices. And this is not just on the case of collusion where you might have agents working together in in unintended ways. But you might also see this, and this is one thing that a colleague of mine is currently working on, where you get this kind of distributed misuse. So there have been various kind of kind of experiments now that have shown that you take some task with some, like, you know, dangerous task, like some constructing some cyber exploit that you might not want some model to do. And if you try and get Claude to do it, it's gonna refuse. You try and get to do it, it's gonna refuse. But if you are able to break down the problem or get an agent to break it down for you, say, like, some open source model where you fine tune safeguards away, to break it down into these individual subtasks, you can just borrow from APIs here, APIs there, etcetera, etcetera. And you can just reassemble all the pieces of the puzzle back together in order to be able to conduct exactly the same kind of dangerous attack that you would have done in before, which is obviously not what we want. And here is, like, really kind of a collective action problem because no individual model deployer is kind of, like, on the hook for this in quite the same way that they would be had I just, like, got Claude to do this for me or got GPT to do this for me. But that's a that's a real challenge. And the moment, my understanding is that the labs do not have any sort of, like, info sharing regime in place where if I see something slightly suspicious over there and you see something slightly suspicious over here, then we can actually kind of join the dots and and do that. I mean, obviously, there are lots of incentives and not to mention kind of, like, antitrust law and these sorts of things that might prevent the labs from kind of sharing this kind of information. But I do think that it could be something like that could be necessary if we're gonna head off some of those challenges.
[25:42] Nathan Labenz: Lewis signed off a few minutes later. And from there, it was the two hosts. Before moving on, I had a request for the labs about the information sharing agreements Lewis had described.
[25:55] Nathan Labenz: I would love to see OpenAI and Anthropic pioneer some of those kinds of agreements that he has been talking about. And they could do that with the independent auditors as kind of the people that get the sort of structured access, the, you know, the sort of private transparency. It could be each other. I think those companies really need to lead in this way. They need to demonstrate that AI can create new institutions that work and not just, you know, swarm and overwhelm existing institutions that, can't keep up with the pace.
Sponsor
[26:36]Claude: Claude is the AI collaborator for problem solvers, helping with writing, coding, financial models, strategy, and more. Get started with Claude and explore Claude Pro at https://claude.ai/tcr
Main Episode
[28:12] Nathan Labenz: Part two, agents in the wild. On September 8, Meta had launched Muse, a personal agent that browses and buys on your behalf. Less than two weeks later, Amazon blocked it from shopping on its site, saying the agent hid its identity and posed privacy and security risks. Prakash and I took that up in a closing segment, just the two of us. His case, whoever owns the customer relationship makes the money. And if agents become the interface, Amazon becomes a supplier to them and loses its margin.
[28:46] Nathan Labenz: Still have a hard time seeing I mean, I I agree that it it does create all kinds of new issues for Amazon, and, you know, they are gonna need to be sharp on this. But, like, I still just don't see why you would rush to ban. You First of all, it's gonna be a small percentage for the time being. I would think there'd be a lot to learn from allowing Muse agents to come shop for a while. Right? Like, I I think one thing that's always always kind of confuses me is why people don't keep the option value open longer. I've always thought this about the chip ban, the export controls. You know? It's like Yeah. If we're going super exponential in chip build out, then anytime you choose to pull the trigger on that still has, like, the vast majority of chips in the future. You know? When we look back on today from 2030 perspective, we'll feel like, well, there weren't that many chips in 2026. And so the ban that went from, like, 2022 to 2026 with whatever, you know, fits and starts and enforcement gaps it had is gonna be kind of inconsequential compared to one that was started today and runs the next four years or even starts next year and runs, you know, for a few years after that. So I don't quite get why they wouldn't let even if it is gonna be sort of an exponential rise in news agents shopping and even if that does kind of threaten their advertising business and, you know, who knows what other issues they might have. But, like, that's kind of the point, I think, at you know, at the beginning, I would first just not wanna turn customers away. And I would also wanna learn what I stand to learn from having these agents running amok, while it's still kind of an early adopter phenomenon. It just seems that they would learn so much from having these these traces. Right? Even if they just wanna build I mean, everybody's talking about distilling and, you know, how to what what are the allowable and not allowable ways to get data. Like, one really easy way to get data would be, like, the agents come use your platform and they see how you you see how you see you see how they use it, and then you can do, you know, all kinds of things, I would think, downstream of that. It just feels shortsighted, but this just doesn't add up to me still. I don't get why you would wanna
[30:57] Nathan Labenz: You are. You are.
[30:58] Nathan Labenz: Wait till it gets to 5% and then make a move. Yeah. The other thing that I do
[31:02] Nathan Labenz: wonder
[31:02] Nathan Labenz: about in terms of,
[31:03] Nathan Labenz: like,
[31:04] Nathan Labenz: what equilibrium are we gonna eventually find ourselves in? I was looking into Cloudflare, and they also have kind of an interesting thing where they're they're starting to ban. But, like, again, it just doesn't feel like the right solution to me because what I then do in response is I have the agent use my browser with all of my credentials.
[31:25] Prakash: Yeah.
[31:25] Nathan Labenz: And you could Good resource. Maybe, you know, detect when, like, an agent is using the browser based on their, you know, sort of bot like behavior. But then you're really setting up an arms race where it's like, this sort of feels analogous to putting too much pressure on the chain of thought. Right? Like, I don't want to put so much pressure on agents that people are out there devising ways to make them indistinguishable from human. I suspect that if people really work hard on that, they can probably succeed or at least often enough, you know, that it will be an issue. And it just seems like it it's much better to have a lane for agents where it's like, this is how they can work. We're not gonna try to block them, but we'll kind of segregate them maybe so that we don't create this arms race. Because I I do I really do think we're headed for a world right now where, you know, probably Meta doesn't do it, but, like, somebody's gonna figure out a way to make agents work on Amazon. You know, somebody's gonna figure out a way to make agents look human enough that they don't get blocked by Cloudflare. And it's also probably works for the user. I mean, I I have this problem all the time where I like, Claude wants to do something in the browser, and this is the same thing happens with OpenClaw, Astra, whatever. They get stuck on when they open up their own window, they don't have sessions. I've given them, like, a lot of passwords where they can log in, but not all. And sometimes they're blocked. And, you know, then if there's, like, a verification flow or whatever, a lot of times they can't do it. And so the fallback just ends up being, okay. Just do it on my browser. Forget it. You know? I I wanted the security. I wanted the separation, but they're making it too difficult. So fine. Just use my browser. Then I'm logged in. You have all my credentials. You can just do whatever. That's not great for me. It's it's not great, I don't think, for anybody. So I do feel like we need a much better solution for all this than just kind of ban and hope for the best. I don't see that holding, and it feels like it also creates a lot of problematic incentives.
[33:27] Nathan Labenz: Part three, the organizations that do not exist yet. Max Nodeau funds technical AI safety research at Coefficient Giving, the funder formerly known as Open Philanthropy. In July, it made its biggest grant of the year, a $160,000,000 to Resolution, the new alignment lab cofounded by Jeffrey Irving. And this month, it launched Project Tailwind, an open call for people to found new safety organizations with checks from $200,000 to $200,000,000. I asked Max what he wants built.
[34:02] Nathan Labenz: One comment that you made on the eighty thousand hours podcast that I thought was quite interesting was that you're excited about funding new independent auditing,
[34:13] Prakash: investigation,
[34:14] Nathan Labenz: verification orgs. What are you looking for when you think about those orgs?
[34:19] Max Nadeau: Okay. So I would about a couple different categories of work that I think is very promising and and important for the world right now. One is on assessing the safety either of of, you know, models or of, like, whole AI companies. And there's there's, like, a lot of different ways of taking the or ways of operationalizing the goal of doing assessments and evaluation of safety. And, you know, we're we're excited about all of them, basically. Like, you know, I'm I'm very excited about more work on alignment red teaming, you know, that that tries to help us better anticipate the ways in which AIs will misbehave before that actually happens. And so so and that's an example of something that you can do kind of, like, at the model level. Like, you can you can do this on an open source model even, and, that might be easier and and more valuable. And then I'm also very excited about about people who are doing more, like, process level, you know, safety assessment. Like, I I think one common reaction that I heard a lot of the face incident is that, like, something went wrong at OpenAI at, like, a process level. And and in particular, it seems bad that, like, the the information about some of these incidents was known to researchers at OpenAI and, like, didn't make its way through the organization to the leadership for for weeks or months or something like that. And then another category of of work which you touched on, which is which is overlapping but somewhat distinct is is around better evidence generation. So it's just really hard for even the most informed people right now to understand, okay, what is the state of AI capabilities? And what is the state of AI alignment? And what is the state of, like, these AI systems in general? Like, well, how do they behave? What are their personalities? You know, in what ways is it is it accurate to anthropomorphize them? In what ways is it inaccurate? Like, you know, people like, we we just really don't have a good science of, like, these systems that we built, how they behave, how they're going to act in new circumstances, what their motivations are if that's an accurate way of talking about them. So better evidence generation, like, can look like a lot of things. It can look like just doing science on these systems to understand them better. It can also look like doing the sort of, like, incident detection, incident reporting work that that Nightingale and and other organizations like that have done. I I don't know if you guys saw these the the these reports about the the German wiki that the OpenAI agents took over. So that was work done by an independent organization that was just crawling the Internet, looking for evidence of AI incidents. And that's work that he's doing because, like, these these sorts of incidents can happen and then just go unnoticed unless there are people who are actually, like, trying to look into them and and present that information to the world and understand the story behind it. So evidence generation is also a diverse and important area of work too.
[36:59] Prakash: To to to what extent do you think organizations like Accenture often function as box checkers rather than investigators. There's a there's a there's a clear difference between an auditor that goes in to kind of, you know, make sure that processes were followed even if, you know, those processes didn't yield the outcome that you are act that you actually desire, an investigator who's there to kinda like, I'm going to find this thing. Right? So so so what is what is the differentiation between those two? And do you think, like, the Accenture's of the world are gonna end up in the box checker rather than the investigator kinda mode?
[37:37] Max Nadeau: I mean, I I know very little about Accenture, so I I won't speak to them specifically. I I think it is a thing to be concerned about in general, putting aside the specifics of the organizations involved though, that these third party auditing groups will will not have enough access or will not have the incentives to pursue really, like, hard hitting assessments of how safety is going at these AI companies. And I think part of why that is is that, like, in many cases, this is all very voluntary. Right? So, like, to my understanding, the only, like, legally mandated role for third party auditors now is to ensure compliance with the, like, self regulatory RSP style policies that these companies are are mandated to have by state law in California, New York, and Illinois. And so, you know, you can just put whatever you want in those policies, and and they're very big in a lot of cases. And so that doesn't leave much much scope or much work for these auditors to actually test anything in particular. So that's that's as it concerns the the, like, legally required auditing. Now when it comes to what labs might voluntarily agree to, I mean, I think there have been some very encouraging signs in terms of what OpenAI and Anthropic have said recently about the sorts of access and the sorts of questions they want these auditors to answer. Now, like, you know, will that actually happen? I think that depends both on whether how how these companies actually implement the the nice words and also, you know, the the auditing companies, how they react to that. So I guess we'll see.
[39:09] Prakash: So you have awarded to Jeffrey Irving, Jeffrey Irving's team, quite a large award. It's structured, I think, as kind of a compute and other kind of award. It's it's too can you tell us a little bit about the award itself, the structure of the award, and, you know, how you guys came up with that with that structure?
[39:31] Max Nadeau: Yeah. Yeah. Sure. But I think what you said is correct. You know, there's a part of the grant that is for research expenses, the biggest of which is is compute. But, you know, when when I say compute, what I mean is both, like, just renting raw GPUs and then also, like, paying AGI companies for their tokens. And, obviously, that's becoming a giant part of all the grants we make is, you know, people increasingly want to run very large experiments, which require a lot of tokens or a lot of GPUs to run open weights models. And also very excitingly, people are starting to have you know, use a lot of AI labor, and that requires a lot of spending on on tokens as well. And, you know, so so AI safety is sort of a funny discipline in in that you can use tokens both at on on labor and on the subjects of the experiments. And in some cases, you're doing both. Like, you're spending a lot of tokens to have some AIs run experiments on other AIs. And so that that creates these these high compute budgets. So so compute, you know, being being the biggest expense is also something where the the actual spending on that has, like, varies by orders of magnitude over time and across grantees. And so sometimes the way we structure these these awards is that we want to be very, very generous with compute and, like, error on the on the the big side. And so we just want to, like, give that as, like, a lump sum that, like, you know, like, this is the money that can be used for for research expenses and not other things. And and that way, we are just very comfortable with the uncertainty that, like, you know, maybe they only end up spending a tenth of it, and that's fine. And, you know, maybe they spend all of it, and then they come back and we'll give them more.
[41:08] Prakash: And Jeffrey Irving has said that superintelligence might arrive in two or three years, and and he proposes that we slow down because that's too quick. Rather than, you know, discuss, like, his views, how do you manage the uncertainty of the time available, and how does that affect what you fund and how urgently you fund something?
[41:32] Max Nadeau: I I mean, I think the boring answer is we just have a portfolio approach. Like, we fund some stuff that is only gonna pay out on really long timelines and other stuff that if it's valuable, will probably valuable be valuable soon. I I think that, like, this question of of timelines to AGI or to ASI, in some ways, matters less for research prioritization than one might initially think. Like, I I think in some ways, prototypical example of a, like, long timelines bet is funding work that, you know, is, like, very theoretical or, like, ambitious or, you know, you you look at it and you're like, oh, this is gonna take a decade to pay off just because it's, like, really, you know, start like like, some work in back in Mechanter or Paul Christiano's work at ARC or, like, these sorts of things where it's like, okay. We're just gonna need to to make a lot of scientific or mathematical progress if this is gonna actually bear fruit in terms of methods that we can apply in reality. And so that sort of feels like work that is maybe implicitly a bet on on longer timelines. But I think that actually isn't really true. Like, you know, if we are going to have powerful AI soon, then presumably, we're gonna have an opportunity also very soon to make use of gobs and gobs of AI labor. And and especially on work that's more mathematical, we're gonna be able to get, you know, five years of progress or ten years of progress in one. And so the sort of work that could be it would take humans, like, a long timelines amount of time to complete might actually be possible a lot sooner under this hypothetical that we're gonna have super intelligence or AGI or something very soon. You know, so long as we actually, like, make use of those AIs and take the opportunity to make a lot of progress very quickly.
[43:15] Nathan Labenz: What would you say are kind of the most speculative things that you included in the project tailwind that have request for projects?
[43:26] Max Nadeau: I think that what we included in the list was, like, a little bit coarse grained, but we have some stubs on the tailwind list about new research centers producing or pursuing more ambitious, more principled bets on alignment. And that is that is something that we've spent a bunch of money on this year and that we are excited about. You know? So so work like ARC, Paul Christiano's org, who is not a CG grantee, but, like, that is a you know, we funded other people working on the ARC agenda and, you know, and working on stuff that is very similar to it because we think that's, like, a valuable Hail Mary shot to to include in the portfolio. And also, I mean, resolution, like, our biggest grant this year is is very much a bet on, like, a kind of crazy thing that has never worked yet, which is, like, maybe we're gonna have some more principled theory of AI that's going to help us have, like, stronger reasons to believe in the alignment of our systems. So that's that's something that we are very excited about, and I think is, like, very weird and out there. And I think a lot of people have very understandable skepticism of people who are, like, who are familiar with what has and hasn't worked in machine learning over the last decade and recognize that there has been you know, all along, there have been people in machine learning who wanted to have to apply theory more to to AI capabilities research. And just like that has not really worked very much compared to just, like, blind groping around and trying stuff, doing trial and error. You know, it's it's like Noam Shazir says, like, the the success of these methods we attribute, like, all else to divine benevolence. Like, like, that's that is the the ethos that has has actually produced progress in AI capabilities. And so I think lot a of people take the lesson from that that trying to do more more theoretical or more principled work on AI alignment is is pretty doomed also. And I I think that's a totally valid reason for skepticism, but we think the upside is is worth it despite that. And so we're, you know, taking that weird bet.
[45:32] Nathan Labenz: Then the people. I put the current conventional wisdom to Max.
[45:37] Nathan Labenz: There's plenty of money. It's all about talent. Obviously, Silicon Valley founder types are, like, one profile that you would want to recruit. Are there other pools of talent that you've identified as kind of strategic priorities?
[45:51] Max Nadeau: Just to be clear, like, the for the things that that CG is is supporting and especially for the things in the tailwind list, money is not the bottleneck. The talent is there are other things within AI safety where where money is definitely a very big bottleneck. In some ways, the the most valuable profile for of talent is someone who combines the the virtues of a founder with also the the thoughtfulness and the the inclination towards and comfort with speculative thinking and futuristic questions that that is more common in the AI safety world. And, like, that's sort of, like, taking it seriously, no bullshit attitude to asking questions about what's gonna happen in the future of AI and, like, really making predictions and, you know, testing your hypotheses and and updating over time. That's, like, a really, really valuable skill to us because it allows people to get going years earlier on on projects before other people can see the value of them. And, you know, an example I give of this is is Meter who before other people were thinking seriously about AI capabilities and and the impacts of them and and how to evaluate them, Meter was very prescient in coming up with a benchmarking approach in time horizons that just worked way better and aged way longer than than other people. And that's because they were they were really, like, trying to actually take seriously what was going to happen in the future with AI and then work back from that to, like, what sort of work they should be doing now that that would age well and prepare for that. So that that inclination is is really, really valuable to us. And, it you know, it's it's useful in basically all profiles. It's useful both in a person who might found an org, but also in a person joining the existing orgs. Like, that's something that existing groups in the AI safety space value a lot because it's it's a big part of how they do their work.
[47:39] Nathan Labenz: Can you do one quick double click on where money still is a bottleneck?
[47:44] Max Nadeau: I think the biggest thing is just, like, in for organizations and in areas that that CG doesn't operate in, money is a big bottleneck. So, like, you know, if if you find some person who you think is doing great work, and for some reason, they they either aren't a good fit for CG or they applied to CG and we weren't able to fund them or we fucked up and decided not to fund them, I think there's there's just still alpha in that. Especially because, like, as we move into making much bigger grants, that can sometimes come at the expense of being, like, you know, covering all of our bases of, like, all the teeny little uses of money that can be that can be valuable. And I think that's the right decision for us to make, but it does leave some things that are unfunded.
[48:32] Nathan Labenz: Prakash should raise the prospect of an anthropic public offering and his expectation that some of the proceeds would be distributed through coefficient giving.
[48:42] Prakash: So what what is your preparation process for handing out larger checks? Like, what is your pathway for the next year, say?
[48:50] Max Nadeau: Yeah. I mean, I I have no idea what's what's gonna happen with the AI company money. In terms of, like, what the preparation actually looks like, I mean, I think the biggest thing is seeding organizations now that can then be ready to absorb lots and lots of money. And so, like, making small grants to new groups and then, like, giving them a little bit of runway to prove themselves or not prove themselves and explode and, you know, like, that's that's the biggest thing we're doing. And and we would love to create opportunities for you know, we we would love to create basically groups of people that can say credibly to big donors. We have people. We have a project ready to go. We're just, waiting for you to sign the check. Well, you know, we've got this, like, shovel ready project, and it's gonna take a billion dollars, and it's gonna solve alignment. And, you know, it's just a matter of, like, giving us the money. And that's, like, really really not the situation right now in that there are just, like there there are just, like, not that many people in this space. And so that's that's a big thing that we're trying to fix.
[49:52] Prakash: And is that and is that about funding people enough so that they can work on things at small scale and such that they can get a much bigger compute check at some at some point to scale that? Is that is that the kind of idea?
[50:06] Max Nadeau: That's the that's the main thing I have in mind. I think the same logic also applies to growing the headcount of their organization. Like, you know, having a group in place that has, like, a great plan that, you know, where they can then absorb, like, a lot of software engineers or ML researchers or or data labelers or, you know, or some other type of person. But but yeah. I mean, I think compute, like you said, is is the main version of that. And and especially, hopefully, AI labor. Right? You know, like, it'd be really great if there were a bunch of, like, Wellscope problems that these organizations had where they were just waiting on lots of AI labor to pour into them. Now we'll see if we can make that come true, and I I think there are reasons to be skeptical that that will happen, but it'd be nice if it did.
[50:51] Nathan Labenz: Part four, the price of compute. Wayne Nelms is cofounder and chief technology officer of Orn, which publishes a price index for GPU compute built from cleared rental transactions rather than list prices. The Intercontinental Exchange has announced plans to list futures on it pending regulatory approval. Prakash asked him to follow a real customer. Who lends against the chips, and who loses money when prices move?
[51:17] Wayne Nelms: So generally speaking, you have a bank or some sort of financer that will lend money against the GPU asset or a host of AI infrastructure assets. What they hope is that they will make their money back over time with interest as a result of the operating activity of a Neo Cloud. So, for example, these are companies like CoreWeave, Crozone, Nebius, Lambda, right, etcetera. However, these business models themselves are contingent on really two things. Right? Continued AI demands and specifically, a ever increasing or at least stable compute pricing, estimate. Right? These GPUs are financed in kind of a variety of ways, including straight line depreciation, which we can get all into. So the way Oren provides value is that you can't really hedge away this risk without having a benchmark to hedge and having a benchmark to trade. So we created that benchmark, and we are actually leading, the transactions on that benchmark for these institutions.
[52:23] Prakash: And how many how many transactions would you get in, like, a month which are which form the index?
[52:29] Wayne Nelms: Yeah. So we collect over a thousand in transactions a day per index. And so right now, we have five public indices. So, you know, it's 150,000, somewhere in that range probably per month. We have a whole host of Neo Cloud partners that we work with effectively that are contributing data to our index every single day. And the question is always, you know, why do people feel the need to show that information to display that? It's because they benefit. Right? A lot of these players are looking for cheaper financing, and cheaper financing comes when their financer has more certainty about the future and can hedge that risk in the future. And so it's actually this nice flywheel where you get more data, you can publish a better index, the financers feel more comfortable and are able to finance more
[53:18] Nathan Labenz: GPUs. I ask how fungible the underlying asset really is, whether an hour on an h 100 from one provider is the same thing as an hour from another.
[53:28] Wayne Nelms: I think it's certainly obvious when you talk to anyone deploying infrastructure or buying that an h 100 or a b 300 is very different depending on the OEM, the supplier, the actual everything about the hardware can be very granular and very different from a engineering point of view. However, our take has always been that there needs to be a benchmark that allows some sort of abstraction for this asset class. I think for us, what's made the job much easier is that, largely speaking, we are looking at one chip producer, which is NVIDIA, which dominates the market, has huge market share. When we look at the NVIDIA moat, and we can get to this later as well, but, you know, you have the hardware, you have the software superiority. But I think the biggest thing that NVIDIA has in terms of a a moat over other competitors is the financing landscape. When you're looking to finance your new cloud and you're thinking about what ships to put into your site, it's just materially better better to have NVIDIA GPUs. Just easier to underwrite. And I think what we see as this market problem is how do we allow every single person looking to deploy compute infrastructure this sort of, almost a protection, for their lenders. We, right now, only reference, NVIDIA reference architecture with InfiniBand, etcetera, etcetera. There are certain steps that we're taking to standardize the compute that we're tracking. However, I will say, like, we are tracking a subset of the market, which might inherently have differences amongst those participants.
[55:09] Prakash: In some commodity markets, you know, there's this idea of baseload where, you know, you have a large amount of baseload, and then you have this idea of, you know, load that gets spun up on on demand. But how do you differentiate the pricing between the kind of baseload versus, you know, the swing capacity? Because I imagine the the the more month to month transaction by the swing capacity and not the baseload itself.
[55:36] Wayne Nelms: Yeah. I love this question. So, yes, you're exactly right. A lot of the capacity is long term PPA style contracts. And I think it's really interesting to think why that is. Right? Like, we're in this situation where AI capacity is so scarce. I think the model providers so OpenAI and Thropic have seen this coming forever at this point. And partly they're to blame for the lack of supply that's currently out there. But if you think about it for them, it's an arms race.
[56:05] Nathan Labenz: Right?
[56:05] Wayne Nelms: Compute capacity planning is an arms race. How much capacity can I lock up over the next few months? So when I need to train the next model, I have enough, and I don't have to go out and looking for more. It also, of course, is naturally a zero sum game where whatever you buy, your competitor can't. So, actually, the way we like to think about the baseload and swing slash, you know, day ahead market structure is in compute, people are buying to fill the peak demand and then selling to satisfy the troughs. Where in power, you are buying to satisfy the troughs, these these long term contracts, and then, buying excess to satisfy the peaks. So there's there's actually robust on demand market that will continue to grow as these training labs are selling back excess as inference providers are selling back excess. It's this reallocation of compute question rather than purely a hedging question.
[57:05] Nathan Labenz: How elastic is the market? You know, like, obviously, we have in an extreme example, Anthropic reportedly has paid up, I think, even, like, a multiple
[57:16] Nathan Labenz: of
[57:17] Nathan Labenz: Yes. Kind of standard market prices to rent at high scale from x AI. How would you sketch out the curve from you know, if I want one h one hundred hour right now, I pay x. What if I want 10,000,000? How much does the price change as my scale as a buyer changes?
[57:38] Wayne Nelms: Yeah. So we like to think about this in three axes. Right? So you have price on one axis. You have duration. So length of contract on one axis, then you have quantity on another one. Generally speaking, as you increase the length of the contracts, the contract per hour value goes down. Right? You're the the logic here is you're buying bulk. Right? You're allowing someone to finance a data center because you're buying ten years of for five years of capacity from them. Generally speaking, when you increase the number of GPUs you're renting, the pricing goes up, which is almost anti this exact logic, which is, you know, I'm buying bulk. Why am I paying more? And it's purely because the number of suppliers that can sell all that capacity at once is very small. You know, you might be able to get a few nodes access here and there. But when you're looking for 10,000, a 100,000 GPUs that are all interconnected or at some level, that is very, very tough. So, you know, certainly, these anthropic space or SpaceX AI deals are few and far between. They're very much OTC trades, I would say. They are indicative of what we are kind of seeing as a general trend across even smaller markets.
[58:51] Nathan Labenz: It is striking that, like, you look at the graphs, and they are all trending up. So that is seemingly in significant tension with the idea that the longer you buy, the lower the prices. Like, why are the are the sellers not AGI pilled enough to expect that they'll command a higher price for the same chips next year than they are right now?
[59:15] Wayne Nelms: Look. Look. I think the the sellers of compute are the most AGI people because they have entered a business where they are selling a at least a resource that hopefully rises in cost in the future. But for a lot of sellers, it's not because they don't want to. I think there's a lot of suppliers of compute that are experimenting with shorter contracts because they know, you know, why am I locking in a four year deal if in two months that capacity is worth even more? The issue here, of course, is risk. It's always this risk reward ratio or or balance that they're trying to play, where even, you know, early financing is contingent on having a long term contract signed. I cannot finance a data center today if I don't already have five years of off take from a credible counterparty signed. And because of that, I can't even enter into the month to month world. So if you actually think about what the futures exchange and the our product enables, it enables an ability for instead of the Neo Cloud locking up a five year contract, they can sell month to month and hedge the future price potential on our plat or through our index. Right? And that is a much potentially a much better risk return kind of math for them and and for their TCO and for their lenders.
[1:00:37] Prakash: So I think the the one thing that, you know, I think I think Nathan has also alluded to is that you have differences in pricing of, you know, near term and much further out, like, term. And you kind of described the use of a lot of the hedging as tools for financiers. And the financiers are often concerned with the far end, not not so much on the near end where they have much more visibility, but the far end where you have this depreciation issue. How do you think, like, the index assists or will assist in the future on this end of end of life kind of depreciation issue, which is, like, four years out, you know, three to four years out.
[1:01:15] Wayne Nelms: Yeah. A 100%. Like you mentioned, a lot of the risk in this space is certainly, long, I would say, long tail risk, in terms of looking at the tenor. And so when you launch these futures, on a futures exchange, right, a lot of the trading can occur in that long tail kind of section of the curve. Right? So these are monthly contracts on ice, to be very clear. You know, we expect to see lots of trading in the front month as people are buying and selling, especially after physical delivery. But even like, I I would expect there to also be a lot of trading in those, like, three to five year spans in terms of that is where all the exposure is. That is where all the risk is. And because there's now features contracts that reference what spot price will look like in three to five years, that is what any bank is exposed to. Right? And to be even more clear for people watching, it's you know, you have a bank underwriting a deal where the GPU life is six years, but the initial contract signed might only be four, right, or it might only be five. And so there's always that, you know, what happens in year four, what happens in year five, what happens in year six when there's not a contract already signed for that GPU. But on your balance sheet, you've claimed that the GPUs have some value. So that is the area where you wanna where you wanna hedge your bets.
[1:02:38] Nathan Labenz: Prakash raised basis risk, the gap between what an index says and what the asset you actually hold is worth. If that gap is steady, a trader can price it in.
[1:02:50] Prakash: But the issue is gonna be if there's nonsystematic basis risk. And one of the issues, know, with GPUs is this introduction of new technology which obsoletes the old and which is what which is what I think the nonsystematic basis risk that I think is, you know, pretty scary for financiers.
[1:03:09] Wayne Nelms: We are tracking GPU specific transactions. So we have a b 300 index. We have a b 200 index. And so, actually, this entire risk of, you know, a financer being worried about obsolescence risk, that is exactly where you can hedge away with this instrument. That is exactly where you can kind of look at how the market is pricing the twelve month or the fifteen month future contract for the b 300. Right? It actually, if anything, allows information to be spread wider, broader, and more evenly Because now you can kind of see perhaps where the market is pricing the announcement of the the announcement of the next chip, right, based on how the curve looks in the future. I think all these things are ways for information to be actually shared more freely rather than for there to be any up obfuscation of this information. And so in effect, I think these markets are a great way to hedge that risk away.
[1:04:05] Nathan Labenz: I also asked how long these chips actually last and whether the a one hundreds from 2020 are near the end of their useful life.
[1:04:16] Wayne Nelms: Yeah. I think the the chip health is a super interesting space to be in and understanding what are the failure modes for the hardware itself. I think, you know, chips will fail all the time for a variety of reasons, but I think what you're pointing to is the useful life conversation of hardware. And this is something that we wrote a paper on just two weeks ago, I think, now about how open source, open way models that are more price elastic will often route to places where the cost per intelligence or cost per workload makes more sense. And oftentimes, those are the older pieces of hardware because they're just less marketable. Right? And so we actually see the useful life extending beyond six years as a general thesis. How much further than that? It's kind of unclear. However, I will say that, you know, what we see in the forward market and the term structure is that there will always be some sort of premium above zero, certainly, and above some terminal value of, like, power. Right? Like, the terminal value of an a 100 is not just the value of the power that powers it because the a 100 will always provide some computing power as a fraction of all computing power in the world. Right? Over time, I think you have a lot of interesting use cases like physical AI, one of many, where these older chips might become more effective. However, today, we see still that these older chips are still being used, still being rented, and have maintained relatively constant demand, in the last twelve months.
[1:05:48] Nathan Labenz: Wayne signed off a few minutes later. In the closing, with just the two of us, Prakash came back to the index with an account of how the biggest compute deals get priced. The figures are his from what he has heard.
[1:06:01] Prakash: What I've heard is that your price of compute is dependent on whether you have financing or not. So if someone needs to use their balance sheet, they get to dictate the terms. So if they don't if if you don't have to use a balance sheet, then all of a sudden, like, you have high pricing. So Elon, for example, Elon had already used his own balance sheet to build out the clusters, and he had the clusters available available. You can take it tomorrow. Right? And he they didn't need to provide him a long term contract. They could do a month to month contract, and they can walk out at any time. Right? So which meant that this is not financeable. A bank cannot come and give Elon a loan based on this contract from Anthropic because a bank wants to see like, if you're gonna pay back in four years, I need to see that you're gonna have cash flows for four years. I need to see that up So so Elon's able to charge high prices because he's using his own credit capacity. He's using unsecured credit at the at the holding company. And he can say, like, look. You you need you you're getting paid $8,090,000,000 dollars a megawatt. You give me the 50 and you're still making money. Right? So that's a good deal. And so Elon got $50,000,000 a megawatt. Right? So I think it changes it starts to change when you need the other guy to give you support. So if you say, like, okay. If Elon said, okay. I'll build it to you. I'll build it for you in a year or year and a half or two years. In order for me to build it to you, you need to give me a five year contract. And if I build it for you, you will definitely buy the compute and pay me. And then I'm gonna take that, and I'm gonna go to another bank, and I'm gonna give it to them. And they're gonna lend me money based on that. Then it becomes like, it's not really like Elon's money. It's really like Anthropic's money anyway. And it's the single credit is gonna be from Anthropic side, and that's really who the bank is relying on for repayment. And so then it becomes a different different different pricing. So then Anthropic will say, okay. How much is it gonna take you to build? How much is your interest rate gonna be? And then I'm gonna give you, let's say, a 20% profit margin on top of that. Right? So that this is what CoreWeave is getting, for example. So if it's gonna cost you, like, $15,000,000 or, you know, $20,000,000 to build, they'll be like, alright. 24, $25,000,000 per year. Right? And so that's the that's the key difference is whether you have the money or not upfront. And those deals don't hit the index pricing. The index pricing is based on willingness to trade. So these are all tradable prices, and that stuff doesn't get traded. And so the people transacting on this index are the lowest peers. They're not Facebook. Facebook can afford to pay, like, you know, a $100,100,000,000 dollars a megawatt. They're not gonna be trading on their own index because, you know, the own index is for people who are willing to sell compute at, you know, like, 10,000,000 or $15,000,000 a megawatt. They're willing to sell the compute. Facebook isn't gonna sell you compute at, like, 15,000,000. Like, they're willing to buy it 50. Right? So so these are, in some sense, the lowest pricings, but it's immediately available.
[1:09:15] Nathan Labenz: Part five, what the machines are actually doing. Nick Gillian is co founder and chief technology officer of Archetype AI and before that led machine learning research on Google's solely radar sensor. Archetype's model, Newton, is a foundation model for sensor data. It takes raw streams from radar, vibration, electrical current, and cameras and reports what a machine or a worksite is actually doing. I asked him to walk through Newton from the bottom up.
[1:09:45] Prakash: Let's
[1:09:46] Nathan Labenz: start with the data. How much data is out there that you were able to just collect on your own? And then I imagine you must have had to establish a sort of partner network to bring in a bunch of data. Tell me about the data process.
[1:10:00] Nick Gillian: Yeah. So we're we're getting close to a billion hours now of of physical AI data that we've been able to script and and gather and collate. In many cases, if you look on the if you look on the web and you look at the different types of physical data that they're that's there, it's it's very sparse. There's very little multimodal data. It's not synced well. It doesn't cover things well. For most techniques, particularly classic supervised techniques where you're basically building one model and it has to have to have a fixed number of sensors and they all have to be captured with the same sample rate, etcetera, that's a massive problem. We've been trying to flip that on its head and say, if we have all of this physical data, which is captured in different sensors and distance different systems with different contexts, how do we actually build the most advanced self supervised techniques that can pull in that sensor data and extract from it in the same way LLMs have done with language. We are able to then join that and mix that in with a small amount of data we can capture from from partners and and and even in some cases, like, that our customers can give us to be able to really extract and understand and map that to physical language and map that to physical control. Right? So there's there's quite a bit of resampling and, you know, really understanding, like, missing values and sensors, is is is that because the sensor had an issue? Or as actually the machine itself is connected to, that's actually a feature that helps explain that machine is about to break. Right? The reason the sensor is giving you all these nines nines is not because the sensor is broken. It's actually actually a feature of the machine, right, in those cases.
[1:11:36] Nathan Labenz: So the area that I've studied best that has done something kind of similar has been at the intersection of language and just images, you know, photographs.
[1:11:47] Nick Gillian: Yeah.
[1:11:48] Nathan Labenz: And we've seen kind of if I had to borrow some of their terminology and map it onto what I understand you're saying, this would be sort of a late fusion approach where you have a language model, you then do, like, a pre trained from the ground up encoder, specialist encoder for these different modalities, and then you're doing some sort of, like, cross attention between those. That's my understanding of, like, you know, at a high level, obviously, of what the image folks have done. How much have you been able to follow in their path? How much, you know, of of the techniques you've seen them develop have worked for you, and how much have you had to go back to the drawing board and reinvent?
[1:12:25] Nick Gillian: We've definitely tried these approaches. We've had some success with them with, you know, other types of sensors, radar and time series and so forth. But the big the big issue that we that we have found is twofold. Number one, there's extreme amount of language sensor pairs on the Internet that can easily be scraped and built into these very large datasets where you have this pairing between an image and a description of that image. Right? This does not exist for radars. It does not exist for time series sensors. It does not exist for all the complex sister systems that we go and and talk to our customers with. Right? So we have had to look at different approaches of how do you, for example, take time series and actually align that with human language. Right? So that's a large part of our of of our research work is how to actually solve this sensor alignment problem. And there's also the case there where with a single image, you may actually be able to get the full context, for example, of the scene and then align that with language. But, actually, with the time series, you need actually need to look at a a temporal window and, you know, what is a single latent mapping to a single word really mean in those cases. Right? It's it's not like you have a time series signal that says dog and you can align that with the keyword dog and so forth. Right? So there's quite a bit of of work we need to we need to do there to solve this alignment problem of how do you generalize this across many sensors. And that's the key kind of components we're building inside inside of our models, which makes them different from, you know, a standard VLM or or other techniques. Right? Because we simply don't have that corpus of of, say, time series and language pairs to kind of apply that previous recipe you mentioned. Right?
[1:14:08] Nathan Labenz: Prakash asked about one deployment, a project with Kojima, the Japanese construction company, and what the inputs and outputs of the model look like there.
[1:14:18] Nick Gillian: In that case, you know, we're working with Kajima, this very, established, Japanese construction company. They they build, you know, islands. They move mountains. The the the project we were working with them on was literally taking a river and moving the river, right, so that it wouldn't flood. Right? So it was a five year long project. They had sensors the whole way down the river as they kind of moved down the river, and they're dredging it, and they're widening it and, you know, they're they're basically mitigating the risk of flooding. Right? And the big problem they wanted to solve is the the their company is outsourcing and subcontracting a lot of the work to a lot of other sub teams, and they want to basically measure to make sure that those teams if they have to go and dredge the river for five hours, did they actually do five hours of dredging that day? Or, actually, they only did two because maybe the the the the the the the digger was in maintenance mode that day, for example, and they couldn't do it or because of weather conditions and others. So for for for that project, basically, the the input to to to to Newton, our our world model, was the multiple cameras around the site, typically three or four cameras looking at a specific section of the river. And when you look at those cameras, the the basically, you have a barge that this boat that's floating out on the river. On top of it, you have a excavator digger. And then on top of that, you at the end of it, it has a end effector that is either drilling or it has the bucket on it where it's actually dredging. Right? They want to know how many hours, essentially, that excavator is actually dredging the river. Not drilling the river, not a crane, not other things. Right? So, essentially, the the input to the system is the sensor cameras, the river sensor hydrometer sensors that basically say how high the river is, And then at different geo locations along that, I think it was multiple kilometer site, where the actual work was supposed to be happening at at that day and what the weather was that day. So we basically have a bunch of cameras. We have a bunch of time service sensors with weather and river river information. It's fed into the model. And the model is essentially outputting what looks like a Gantt chart. Right? It's actually outputting the state of that team. Where is the barge? Is it parked on the side of the the river? Is it actually moving, into position, which can take an hour sometimes to do that? Is it now actually in position and they're starting to dredge? Or is it just there and stationary and nothing's happening? Right? And then they can start to map these work charts over multiple years of data and start to look at these bigger patterns that a human can never spot. Like, one of the patterns they they they find, for example, is obviously, if there's a storm, the work on the day of that storm goes down. That's kind of, you know everyone would know that. Right? But one of the key things they find is multiple days after the storm when the weather had cleaned up, the work productivity of the team was still very low. And the reason for that being is the the the the the river water coming down from the mountains, it takes several days for it to actually reach the end of the river and go to the ocean, is where the site was. So So there was a lot of debris and a lot of churn coming through the river multiple days after these larger storms, which is actually slowing the teams down quite a bit. They'd never spotted those patterns before because they could never look at these aggregate patterns. But Newton was able to kind of spot those patterns for them. Right?
[1:17:45] Prakash: How much does Newton generalize in the sense that when you go into a new site
[1:17:52] Nathan Labenz: Mhmm.
[1:17:52] Prakash: Do you have to take all of the data and kind of fine tune the model once
[1:17:56] Nick Gillian: Yeah. So we are building the model so that out of the box, you can plug in different sensors and different configurations, and the model can automatically adapt to those, you know, as much zero shot as we can. For a subset of users, they can directly do that. The the typical analogy I give customers is if I if you showed this task to, like, an average person from the street, would they be able to detect the safety violation? Or would they be able to detect that looks like a strange signal and anomaly? If the answer is normally yes, normally, means that Newton can do it out of the box. It doesn't need, you know, a specific, you know, master technicians level of expertise yet to be able to spot that. Right? For those cases where maybe the piece of equipment the user is using is very specific to, you know, that factory, for instance, or they have certain parts that, you know, they they have very internal and, you know, unique names for the for those processes or parts. That's typically where fine tuning or being able to bring customer's knowledge base and actually being able to inject that security into Newton really helps. So in that case is what our our platform allows customers to take our model, run it directly on the customer's infrastructure so no data is leaking out of their their networks, either on on-site, on prem, or in their secured clouds, and then be able to take their own data and actually do some amount of fine tuning to be able to customize the models. And, typically, with customers we've seen, you know, on the order of a few thousand samples is enough to be able to take the general model and then adapt it very quickly to, like, that customer's use case.
[1:19:24] Nathan Labenz: Last question for
[1:19:25] Nick Gillian: me
[1:19:26] Nathan Labenz: is about your vision of superintelligence and how you might relate to the frontier companies as we go forward in time.
[1:19:35] Nick Gillian: Mhmm.
[1:19:36] Nathan Labenz: What I think is, jeez, the reasoning is already getting to, like, super intelligent light, you know, with all these Millennium Prize problem style results happening just through super long chain of thought. Like, who would have believed you could get there? So on the reasoning front, we've kind of come awfully far. What we don't seem to have still is, like, a good intuition for the many domains that matter. And what I sort of expect to happen is something similar to what has happened with images where, you know, it's one thing to sort of put an image into a model and get a caption out. But when you can get the altered image out and it's clear that the model understood both your verbal instructions and also sees the image in a very similar way to the way we intuitively see the image, which is something we are really good at, then it feels like, wow. Okay. It's really got that domain.
[1:20:30] Nathan Labenz: Mhmm.
[1:20:30] Nathan Labenz: So my superintelligence vision is basically do that across a ton of different domains, many of which we don't have native senses for with what you're building being a great example of that. Right? People don't have a native sense for how to integrate 10 signals from 10 different sensors. Do you see a role for you in superintelligence being sort of building the deep intuition for this domain? And does that ultimately get joined with, like, elite reasoners via the form of, like, a bidding war between a couple of frontier companies for your company, for example?
[1:21:06] Nick Gillian: I definitely see the part that we can play in the in a larger superintelligence system. So that's kind of how we are seeing, you know, our our where we play in this kind of bigger superintelligence loop. It's like we're very focused on how to understand the world around us and be able to bring that back to natural language and machines so that these things can talk to each other. If we can do that and other folks are solving digital intelligence, that nicely will connect together very well.
[1:21:34] Nathan Labenz: So just to make sure I understand your answer there, it sounds like you are envisioning a world where the technology that you're building remains sort of a tool call type relationship to a core reasoning model and doesn't so much get, like, upstreamed and thrown into the, you know, GPT seven pretraining mix to spit out one model that does everything in the same latent space.
[1:21:58] Nick Gillian: We're we're thinking of it more on on the line of we're building physical agents that can be deployed out to the ecosystem of the world. Those physical agents then can talk to other agents and be able to, you know, bridge that gap between the digital world and the physical world. How much that's actually done on an agent to agent, you know, post training mechanism? How much of that is pulled into the pretraining? Get to see. Particularly when it comes to those more nuanced sensors like radar and lidar and, you know, infrared and all these things that don't look like human, you know, visible images. That'll be interesting to see how that's brought back into these kind of bigger print, you know, pretraining loops versus just it's more of a an agent that can go and solve that problem and tell you what's happening in in in the physical world in real time and be able to let you, you know, interact with it, for instance.
[1:22:50] Nathan Labenz: Part six. The dish is not the body. Andrey Georgescu is cofounder and chief executive of Vivodyne, which builds robotic labs that grow pieces of human tissue, lung, liver, tumor, and more from primary human cells and tests drugs on them. In August, it opened what it calls a human biological data center, 12 robotic labs that, by the company's count, can grow more than 3,000,000 tissues a year. Prakash had the first question. How how
[1:23:22] Prakash: do you actually, like, figure out, what the dosage is and whether or not, you know, that is actually what the cell would encounter in the body?
[1:23:33] Andre Georgescu: Basically, everything that's been done, preclinically to date, right, outside of testing in animals, you are just dunking, you know, some substrate, whether it be an organoid or you have cells in a petri dish, just basting them with this with this test compound, that you wanna test. And that's exactly the problem is is, like, that's not how transport happens in the human body. Right? And I think it's a perfect example. A lot of cell therapies that are designed for solid tumors, in a dish will completely kill that tumor. And then those same CAR T cells, when they go into a patient, just flow right by the tumor through the blood vessels. They don't even know it's there. Right? And then and and thus is the challenges, modeling the transport of that therapy into the tissue. Do they recognize the tumor? Do they concentrate there? And so in, vivaDyne tissues, even the blood vessels, the the kind of capillary bed in that tissue is self assembled. So these little blood vessels grow within the tissues, and we can perfuse directly through them. And whenever we dose a drug, it is always through the blood vessels that are native to that tissue. And so the transport through the interface of the blood vessels and then the, you know, like like like meat of the tissue, right, all all of its, like parenchymal, its its functional subparts is just like you would find in a normal tissue. The density is the same. Right? And so, that is just as big a part as does the drug molecularly, like, bind to its target and perform this effect as does it even get there.
[1:24:58] Nathan Labenz: Then my question on scaling laws.
[1:25:03] Nathan Labenz: I'm very
[1:25:05] Nathan Labenz: intuitively
[1:25:06] Nathan Labenz: bullish on this just go collect an obscene amount of data on this kind of tissue perturbation paradigm. It makes a ton of sense to me. The skeptical view that I've heard from time to time is you can never have enough data. You know, the the world is too big. Biology is too complex. You can do that, but you just won't be able to generate enough. The the the the
[1:25:29] Nathan Labenz: challenge
[1:25:29] Nathan Labenz: of
[1:25:30] Andre Georgescu: this argument that it's just like this just a, size of dataset thing is that currently, even the state of the art virtual cell models that exist saturate at a very, very small, fraction of the input data that is fed to them. So you can have all the data you want, but their performance saturates after a couple percent. Right? And this is, like, you know, both published and you know, what we hear on the ground, right, in in in these perturbed seq kind of based models. And the reason it gets so difficult is when cells are growing in a dish, which is which is the the the kind of mode for this for this sort of perturbed seq, they are so far removed from all the feedback loops that are natural in people that they're just trying to colonize that piece of plastic. Right? It's like you have a stiff substrate. Normally, cells when they encounter that are like, I have to encapsulate this, and so they are just growing. Their only goal is to proliferate. And so if you knock out genes in those cells that would otherwise affect how they perform their natural function, really, there's no big difference. Right? They're they're still just dividing. And so you end up learning this view that perturbations don't really do much of anything or it's unclear what they do, except if a gene is, like, required for that cell to live, in which case, like, it kills the cell. Right? And so and so it's not there's not even causality to extract from that. And so we we think that the that the that the fundamental bottleneck to having this causal understanding of of biology is that that we need to perturb these tissues in the complex environments that they call home naturally. Right? Almost like, the the the image in my head is we are trying to populate a map and kind of like design a GPS through it. We can say if I wanna get to the state, I need to, you know, like, go down this road and then turn left with with a second perturbation, and it'll bring me to this healthy state.
[1:27:21] Nathan Labenz: I
[1:27:22] Nathan Labenz: asked how you get from a sample received to something placed in the system and growing.
[1:27:28] Andre Georgescu: And when we grow these tissues, you know, a a a really big differentiator for for our approach to doing this, and I think the the, you know, really, like, critical mechanism that you need to to get, like, realism in this, right, is that we use primary cells, which are which are mature differentiated cells in in that that, like, perform our natural organ function in our bodies, as opposed to starting with some stem cell completely absent from all the conditions that would tell it how to how to mature into a into into a differentiated cell, and then, like, sledgehammering it with cytokines to to turn into some cell type that we think is accurate to what happens in the body. Right? We we instead just kind of go to the mature source, the the ground truth. And we we, mix together all of these different cell types that comprise, you know, the lung, for example, a very high density. And within that injected tissue, we see that these cells, begin to reassemble into the native structure of that human tissue, right, with some little herbs and spices that we introduce. But the the, the the, like, responsibility for forming the native structure of human tissue is on those cells that already know how to do it. It's honestly kind of injection molding into a into a, like, into a chamber. Right? So we we, like, shoot in these these, cells at, like, at high density. And then over the course of several days, they they reassemble themselves. They just self organize to make blood vessels to to partition into, you know, the in the case of the lungs, like, airway where there's air and then the parts where there are blood vessels and fibroblasts and immune cells and all of this. So we are we are relying extremely heavily on the cells to, be be able to direct their own self assembly at these very small scales.
[1:29:12] Nathan Labenz: Can you tell us a little bit about the mix of in silico via the foundation model versus in cell in in actual tissue experimentation today and, like, how that Yeah. Ratio is expected to evolve? Because presumably, like, the big value is it's a lot cheaper, faster. Even if you've got all these great wet lab technologies, just plowing things through the foundation model is supposed to be where you get, like, the Yeah. Insane acceleration. Right? But then, of course, you gotta ground it and validate.
[1:29:42] Andre Georgescu: That's exactly right. So so
[1:29:44] Nathan Labenz: so if
[1:29:46] Andre Georgescu: we if we look at it like a reinforcement learning problem, which is, you know, my my actor is this is this, well, it's kind of the experiment. Right? I have this tissue, and I can perturb it in very many different ways. And the goal is I have this objective, which is to make it healthy, and I can instrument the state of that tissue very deeply. So I know, like, the the the value function is very well defined, in in terms of the state of that tissue.
[1:30:11] Nathan Labenz: The the
[1:30:12] Andre Georgescu: question of what should I try next, right, in in the in the large rollout that I can perform, what are the most use useful things to try next that add the most information to to my dataset? So the the physical experiments kind of serve to to, like, to, like, fill gaps in which the model is most uncertain. And the value of the model ultimately, right, like, why even make these foundation models is, the the the combination space, especially when we move into, two target or three target therapies. Right? If we have, just combination therapies, if I take two drugs because it's gonna work much better than than one, is is so huge that you have I mean, you could have like, the the this whole planet could be, you know, vividine systems growing growing human tissues, and you'd still not get anywhere close, right, to the to the just sheer space of things to to test. And so we need some way to narrow down to, you know, to to to kind of just, you know, a couple times, like, the total number of clinical trials in America, the the experiments that that that we should physically run to, to to kind of test our hypotheses. Right? And so and so, basically, these foundation models allow us to predict what happens when we drug this thing and that thing or all these things together or, you know, this thing and then that that thing. And we we kind of optimize and we iterate in that space, and we find, for example, the thousand most likely things that will, you know, move this tissue closer to this healthy state. And then we close the loop by testing it. Right? And we see where that rollout goes. And then we just do that on a loop. Right? So the the the the purpose very fundamentally for us of these models is to tell us what to do next, and then the purpose of this experimentation is to figure out if the thing that we thought we should do next actually does move us closer or or how it deviates so so we can correct.
[1:32:06] Nathan Labenz: Prakash asked about the hardware. The tissues grow on what Vividine calls a tissue disc, and the company recently announced a second generation of it. He asked what changed from the first disk to the second. Andre started with the robotic lab the disks go into.
[1:32:23] Andre Georgescu: Imagine something the size of, like, a wardrobe. Right? So it's it's about about the footprint of a large desk, about eight and a half feet tall. And these are self enclosed labs. So fridge and freezer, and there's a, know, three d scanning microscope in there and, you know, big robot arm that can use all these different tools. We wanna optimize the surface area of those disks for tissues. And so in our in our second form factor, we have an increase anywhere from two to four times the density of tissues that we grow on each disk without any sacrifice to they're actually much larger than the than the, you know, version one of these disks. So we increase the number and the size of these tissues, and we offload a lot of the responsibility for how they're grown to the robotic system around them. Now the the the challenge there and, like, the the kind of work that allowed us to do this is, a lot of work in software on, like, the on the orchestration layer. You think if you had, like, 10,000, experiments run-in parallel, and this is not like I have a, you know, different gradient of dosing for some drug or something. Right? I can do it on a plate. It's I I am changing the collection point of data. I'm changing the interval between dosing with some drug. I have a, you know, a first stage regimen and a second stage, you know, therapy. I'm taking all sorts of different readouts of of of these tissues. It becomes a super, super complicated experiment. And I can define it, like, quickly, like, by by you know, I I I I can explain it out loud and say, well, we wanna test, you know, bone marrow and airway. We wanna dose these drugs in this order. We wanna branch to these concentrations, and I want three d imaging, single cell sequencing, and proteomics. And and and by saying that, I can define the study decently, unambiguously. Right? I mean, like, given the system of tissues that we grow, it's pretty unambiguous what I would wanna test in there. But then the robot has to be able to decompile or to to to to to compile that, kind of zipped representation of the study. It it has to it has to decompress it into this this, compiled list of, like, millions of actions that this robot has to perform. It has to remember that if it took the lid off of a plate, that it should put it back, right, as an example. But if there's multiple trays that it has to handle and they're all out on the on the working deck, maybe it should delid multiple of them in advance. Right? If multiple pieces of plasticware are on the same tray and they're coming out of the freezer to thaw, you wanna make sure you're not thawing something that should not be thawed too early. And so all like, this this this compiler behaves a lot like a compiler in in code, right, for c, where you're managing, like, the prefetching of information for memory and branch prediction, like, all this stuff, right, that that that, thankfully, our operating system and and and, you know, hardware handle for us. But improvements on our orchestration mean that the system itself, the hardware, the automation can do much more of this, and more and more of that space on these tissue disks can be used for actual tissue biology instead of having to integrate, like, the little, you know, helpers on there. Toward the end, Prakash read Andre a grand challenge for the field, which he credited to the annual review of pharmacology and toxicology.
[1:35:38] Prakash: On the annual review of pharmacology and toxicology's grand challenges in modern pharmacology Uh-huh. The innovation, a universal in silico or organized framework that forecasts rare patient specific adverse events with more than ninety five percent accuracy before first in human dosing is one of their target target challenges. Vivodyne is obviously on on the path there. When do you think this challenge will be solved?
[1:36:12] Andre Georgescu: So this is a hot take, but I feel like I feel like, this is almost like, picking the color of my Mars suit currently to me. Right? Which is like, how about we just, like, kinda fix the bigger problems? And and and the big problem by the way, we we we are working on specifically that challenge, right, of of of of of, like, very, very rare outcomes. But a bigger like, far more regularly, we have drugs that would have huge potential in their efficacy in in patients, but that have routes of toxicity that are pretty common. Like, many patients, like, die from it. Right? I mean, many of the cancer drugs that are going to clinic today, the the the risk is not like, oh, man. You know, I've I've cured nine hundred ninety nine people of cancer, but this one patient, like, didn't work so well. You know, it it it is more toxic. The problem is, in all the patients we've dosed, they barely had any improvement in their cancer, but it gave them all of these side effects or killed someone from liver toxicity or something.
[1:37:10] Nathan Labenz: Andre signed off a few minutes later. In the closing, with the two of us, Prakash put the two guests side by side.
[1:37:17] Prakash: I liked how, you had basically two different approaches for data. You had, you know, on on the one side, you had let's take all the messy data. Let's feed it into the into the machine and try to make sense of it. On the other, Andre is like, there's just too much messy data in biology. I'm gonna build my own data collection effort, and I'm gonna monitor every single thing that goes in and comes out so that I know exactly what's going on. And I and I think I think Andre probably has a good chance of solving this thing. Like, you need you need a little bit more GPU power. You need a little bit more imaging. I think the imaging granularity is probably not there yet, but you can definitely see if they manage to scale a little bit more. The the real question is, like, do you need to do 3,000,000,000 samples, 30,000,000,000 samples a year? Do you need to get everyone genetics, or are you gonna get is your scaling gonna be able to get you to a point where at 30 or 300,000,000 or wherever, it just starts to generalize a lot more and you stop you stop needing so many so many different samples?
[1:38:23] Nathan Labenz: Prakash guessed it would be two or three more years before AI gives biologists something they can really use. I took the under.
[1:38:31] Nathan Labenz: I think we're gonna see probably a reversal of the order of, like, utility and Millennium Prize solutions in biology relative to what we've seen in math. The issue with math is that we already had a ton of math that was, like, interesting to mathematicians but not super useful. And so to push the frontiers of math, you kind of had to get into the, like, not useful domain of math. In biology, we have, like, tons of useful stuff that is, like, still not even that well understood. Right? I mean, it's just very it's a far more empirical than theoretical domain, And I would expect a tremendous amount of utility to emerge before we have anything approaching, like, a full, you know, mechanistic or fully causal understanding of what's going on in in biology. And to some degree, you know, obviously, we'd love to have that,
[1:39:21] Nathan Labenz: but
[1:39:22] Nathan Labenz: we don't necessarily care, right, at the level of, like, does this cancer drug work on this patient's tissue? If it works, you don't have to have a mechanism. Right? That's not part of the clinical trial approval process. It's nice to have.
[1:39:37] Nathan Labenz: Not required. Prakash also made the case for specialists. A model like AlphaFold does its one job better than any generalist. He expects that to hold. And the version of superintelligence people fear is a single model that does every job better than the specialists and depends on no one. My answer was the pattern as I see it so far. New kinds of data get pulled upstream into the biggest model and efficient specialists get distilled back out of it.
[1:40:06] Nathan Labenz: But, yeah, put me down for one that has not seen a reason yet to believe that all these modalities don't end up integrated at the frontier. In a way, I wish it wasn't gonna go that way. I do
[1:40:17] Nathan Labenz: think
[1:40:17] Nathan Labenz: safety through narrowness would be really nice. That's why I'm excited about JEV in part, you know, because it's sort of it has a very niche role to play. You could build all kinds of scaffolding around it. I was a big fan of the Eric Drexler reframing superintelligence piece, which he he also known as comprehensive AI services, where you've just got narrow purpose built AIs for all kinds of different niches. My hope for for safe superintelligence is something like that, you know, just based on kind of reading the tea leaves of Ilya's comments. Could we possibly get a model that starts as kind of a stem cell generalist and then matures in a way where it gets really good at its job, but also kinda settles in to that niche in the way that mature cells don't revert and turn into other kinds of cells and that they're kind of limited now. They've lost their pluripotency, so to speak. I love all that stuff. I I will like, I hope it goes that way. For efficiency, it probably will. But at the frontier, I just have not seen anything at all yet that makes me think that the the single biggest best teacher model you could make couldn't just learn it all. And that you know, given what we've seen, that that starts to be a bit of a a scary beast.
[1:41:33] Prakash: No. I I I I agree. It can learn it all, but I think the learning it all at the most efficient compute rate and, like and and inferring at the most efficient compute rate, which is what you alluded to with the distillation. I think that's the that that doesn't happen. Right? So the energy efficiency is is kinda key because the energy efficiency means that you have, like, this diminishing returns curve as an economics for getting bigger. Okay. The doomer scenario is less likely to happen because you will get a more intelligent model, but it's gonna be less efficient to start out with, and it's gonna be more energy consumptive. And that means that it's gonna need, like, more energy expansion to get there, but the other models are gonna be more efficient than and more numerous. But I think the singleton example, I think, is not likely to happen.
[1:42:20] Nathan Labenz: From your lips to god's ears, as always, I think that that's a real constraint for sure. I mean, if you made a you know, let's just get ridiculous. Right? You made a quadrillion parameter model, you certainly would find it a bit slow and a bit expensive, and you couldn't run a million agents at that scale of model today. So I do think there are some practical limits there. But then again, remember what we talked about on Wednesday, there's gonna be more compute installed over the next twelve months than exists currently in the world now.
[1:42:54] Nathan Labenz: Yeah.
[1:42:54] Nathan Labenz: So we are definitely going to test for at least a couple few more years whether just bulking up on compute solves all you know, quote unquote, solves or perhaps creates all these problems.
[1:43:10] Nathan Labenz: That is the week. Tell us what worked and what did not. See you in the morning.
Outro
[1:45:16] If you're finding value in the show, we'd appreciate it if you'd take a moment to share it with friends, post online, write a review on Apple Podcasts or Spotify, or just leave us a comment on YouTube. Of course, we always welcome your feedback, guest and topic suggestions, and sponsorship inquiries either via our website, cognitiverevolution.ai or by DMing me on your favorite social network. The Cognitive Revolution is part of the Turpentine Network, a network of podcasts which is now part of a sixteen z where experts talk technology, business, economics, geopolitics, culture, and more. We're produced by AI Podcasting. If you're looking for podcast production help for everything from the moment you stop recording to the moment your audience starts listening, check them out and see my endorsement at aipodcast.ing. And thank you to everyone who listens for being part of the cognitive revolution.