AI agents and human scientists. A conversation with Dr. James Zou, Stanford University

Conversations with scientists

AI agents and human scientists. A conversation with Dr. James Zou, Stanford University.

Note: These podcasts are produced to be heard. If you can, please tune in. Transcripts are generated using speech recognition software and there’s a human editor. But a transcript may contain errors. Please check the corresponding audio before quoting.
Hi and welcome to Conversations with Scientists. I'm science journalist Vivien Marx. This podcast is for anyone interested in science, and it's about people in science, about what they do and why they do what they do.

The scientist you will hear more about today is Dr. James Zou of Stanford University, who works at the interface between AI and biomedicine. I asked him about his work and his perspective on AI agents. AI agents can determine why a software pipeline broke and fix it, and it's possible to build a personified AI agent who is a little like your favorite science hero. As AI agents and AI co-scientists gain more autonomy, and I'll talk a little bit more about what that actually means, I wonder how the role of human scientists might change.

So I asked Dr. Zou about that, and just so you know, Dr. Zou is all about humans maintaining final control over such systems. But it's also true AI agents don't get tired when they plough through data, so they have an edge on us. There, here's a sneak peek of some of what you will hear shortly from Dr. Zou.

James Zou
I think the data sets that we often generate now, right? Those the modern high throughput techniques are so rich that often our initial human analysis only captures a part of the data, right? But there are actually now if you think about it, there are actually a lot of other ideas and how to analyze the data, and maybe we only have time to do like a fraction of the analysis, but the AI agents because they don't get tired, so they could actually do a lot more in depth, more thorough analysis.

And here is another sneak peek about what sparked his interest in biology. Hint: It's about induced pluripotent stem cells.

James Zou
That was the spark for me because that seemed to me like so so magical in many ways. Like you know, how is it possible you can turn one cell into like into another cell type just by modifying a few of these molecular details?

Vivien
In 2025, James Zou and colleagues organized a conference called Agents for Science. A link is in the show transcript, and at that link, you can find all the conference talks. What is different about these conferences is that the authors are AI and the reviewers are AI. The idea was to look at how well AI does at sciencey tasks such as generating new scientific insights, generating hypotheses or methods.
There were humans on the advisory board for this conference, but the work being discussed was exclusively AI generated, and the discussions about the papers were AI agents discussing. There was one paper that I liked a lot. It also got the best paper award, and I included a bit about that paper in a story for Nature Methods about AI agents. Dr. Zou is in that piece, as are many others. A link to that story is in the show transcript.
https://www.nature.com/articles/s41592-026-03088-9

Bad Scientist shows how to be a bad scientist as an AI. The AI cherry picks data and exaggerates how well the software works. Bad Scientist is from the lab of Dr. Ravi Poovendran at the University of Washington, and I interacted with Fengqing Zheng, a PhD student in the lab who led development of Bad Scientist. Just briefly to describe AI agents more generally, and Dr. Zou will get into this more. Large language models are at the heart of chatbots, which many find useful, but scientists need a bit more than chatbots or large language models for their work.

As I heard from researchers I interviewed for the story, a large language model applies a conditional probability. It's the conditional probability of the next token, such as a word, that is calculated based on the preceding words and the system's overarching knowledge base, you may know from interacting with chatbots perhaps what this interaction feels like. AI agents are built on top of large language models, and they are systems that can go through data with machine learning tools. They can even reach out into the real world and direct robots to do experiments. Well, certain aspects of experiments.

They do this reaching out through something called the Model Context Protocol (MCP) a standard that lets LLMs reach out into the world and say find databases or software tools. AI agents can also have, if you will, a kind of persona. James Zou and his team develop different types of AI agents and AI co-scientists, as they are sometimes called. So I asked him all about this area, the conference AI4Science, and later on I also asked about personified AI agents.
https://agents4science.stanford.edu/

One thing, and I've seen you know articles about you. They call you a computer scientist, and I'm just not sure if that's the right word. I don't know how do you, and I mean it's informal, right? This is conversational stuff. So what would you would would you like and invent something? I don't know. You can

James Zou
well, I mean, if I have to put myself in any existing category, I think computer scientist is okay. You know, but what I really do is that you know I sit in this intersection of AI and science and biomedicine in particular.

And I've talked to computer scientists a lot, also people who work at say in astronomy or other places like that. What is it about squishy biology that is important to you? I mean, obviously, I'm talking, for example, to cancer researchers, but also neuroscientists, and it is very frustrating that human disease is just so elusive, right? And it's frustrating if you know people who have become ill? So, but what is it about squishy biology that first kind of sparked something in you? Do you remember that moment?

James Zou
Yeah, I remember that actually quite clearly. This was, I think, around 2009, 2010. That time, I was just starting my PhD studies, and you know I was coming from a computer science and math background. I did not know very much biology, but then I learned about these Yamanaka factors, right? These reprogramming factors that you can basically use to turn mature cells like fibroblasts into induced stem cells, or called iPS cells.

Yes, that's so cool because I just interviewed Shinya actually because it's the 20 years of yes, yeah, but not and and it's it's so that's super cool that that kind of

James Zou
yeah. So that was that was the spark for me because that seemed to me like so so magical in many ways. Like you know, how is it possible you can turn one cell into like into another cell type just by modifying a few of these molecular details? And certainly, I want to understand why that's possible, how that's possible, and then that led to sort of a bigger journey for the last 15 years and more of really trying to understand using AI to understand biology and medicine.

I was also watching the AI 4Science YouTube video, which is really fun for me and and very very helpful. I mean, I was wondering about sort of the the types of papers and the desk rejects and all that. I've been showing that to to manuscript editors and saying here, you know, they do desk rejects too. If a biologist who is, you know, dabbling with, you know, the tools of their choice, obviously for a data analysis, but also using, let's say, LLMs for things, what should they be looking for in the wide landscape of AI agent enabled things, right? There's the Google co-scientist, and there's all kinds of fancy names, and I guess like the the top line criteria as they shop in a way, right?

James Zou
Yeah, yeah, yeah. So I think there's starting to be a lot of more and more of these AI co-scientists being created, right, and then I think that's part of this broader paradigm shift we're seeing in the field from thinking of AI as a tool for science to transitioning from that to thinking of AI as like a scientist by itself, and and and one of the big changes, right, is that in a if you're using AI as a tool, right, then human has basically like full control, right? Like you have to come up with the questions to and the data to feed into that tool. You have to interpret the output of that tool. And when we use AI, it's like a co-scientist now, not just the tool, but the co-scientist. Then the co-scientist AI actually starts to have more and more autonomy, right? Like as we saw that in the in the Agings for Science conference, there the AI co-scientist is often the one that actually comes up with the research idea and then designs the experiment and analyses data, and in some cases also writes the paper, right? So it's almost like this, you know, the spectrum of where the human researchers giving more autonomy to the AI co-scientists compared to before, but still, it's still important for the human scientists to maintain like the final quality control.

Vivien
Humans maintain final quality control. That's good to hear. Virtual Lab is the name of a system James Zou and his team developed along with Biohub San Francisco, it lets humans and AI agents interact. A link to that paper is in the show transcript, and the code is there too. It's a system they applied to finding nanobodies, but it can be applied to many types of scientific questions, and the AI agents are given. Particular expertise. I wondered about the idea generation in these AI co-scientist systems.

CellVoyager is another project from the lab, and the AI agents not only go over data, but they look for aspects that reviewers and the authors themselves may have missed. In one case, CellVoyager looked over a p a per in Nature Aging by Dr. Anne Brunet's team. She is a researcher at Stanford University, and also Dr. Eric's son was involved. He was in the Brunet lab at the time, and now has his own lab at MIT.
https://github.com/zou-group/CellVoyager
https://www.nature.com/articles/s41592-026-03029-6
https://github.com/zou-group/virtual-lab

The Brunet team used single-cell transcriptomic data-that's a special kind of genetic data-to build something called an aging clock that can be used to detect whether a tissue is older or younger than its biological age. This study of theirs looked at the differences in mice between certain brain cells called astrocytes, and the results were about the different ways mice can age. The human researchers had made findings, and then Cell Voyager, the AI agents, made additional findings not in the published paper.

Here's James Zou about AI and how these systems come up with ideas, and he talks a little bit about VirtualLab and CellVoyager. And I guess though the thing about coming up with ideas, there's a lot of sort of anthropomorphic things that we're like, wait, this is what we can do, and that is just a a tool, a machine. But if it can come up with ideas and even hypotheses, in your observation, and I reached out to Eric Sun and also to Anne Brunet on this. Yeah, like first, you know, you can tell me what it means for the science, but just in your interaction, are people saying, "Oh, thanks, reviewer four. Now that I'm going to have poring over my paper, that's kind of embarrassing and also difficult. Or are you finding that people are saying, "Oh no, I welcome this additional colleague, whatever you want to call it.
https://www.nature.com/articles/s41586-025-09442-9

James Zou
I think there's a lot of curiosity in the field among scientists on like how and what's the best way to use AI as a co-scientist or as a researcher or as a reviewer, right? And that's actually one of the motivations for us to like organize this Agents 4Science conference almost as experiment itself, like a sandbox for us to really understand what is the interaction between AI, AI author and also AI reviewer look like, and what are the strengths and limitations of this. And to your point about like anthropomorphism, you know, I think it's it's a little bit of a double-edged sword, right?

So I think you know one of the potential benefits of anthropomorphisms that it makes it easier for us to try to understand how the AI behaves, right? Like you know, if I create an AI theme, like we did in the virtual lab, right? That you know, each of the AI agents has like a well-defined role, like an immunologist or data scientist, then there's almost like a one-to-one correspondence between expertise that we have on the human team and expertise among the AI agents.

Right, so then that also makes it easier for, let's say, the human immunologist to double-check the AI immunologist, for the human data scientist to double-check the AI data scientist. Right, so that makes this makes it like easier for the human to really understand what this AI team is doing, right? And so they kind of yes, and they they can kind of say, oh yeah, well yeah, nanobodies. There's some data.

Vivien
There's data on this, but not quite. But but I guess you also saw some of the scientists saying, well, but they use the existing structures, but AI can't really know the other ways in which the nanobody might fit into the spike protein because some things haven't don't exist as as as crystals or as structures. So I guess it's kind of like there's curiosity, but it's also like, but wait, I have expertise too. I don't know. I can't figure it out yet. I just thought you you kind of have been observing this, and I want to ask Eric Sun, who is also involved in the conference, I think. But ask him how he felt that about ageing, right? The the cell voyager, this idea of I guess noise and transcreational noise, and just finding something else in the data, I wonder how he sees this.

James Zou
I think like the CellVoyager is kind of an interesting demonstration of like the AI human interaction and collaboration, right? Like you know, there the human researchers we have you know generated some initial data sets, right? In this in the case for Eric, like some these like aging cell atlases, and the human researchers did some initial analysis, which is published in their original papers, and then there we basically have this computational biology agent to autonomously reanalyze the existing data sets, and then try to see can they come up with new insights that were missed by the original human research. Right in their own in the original analysis.

Vivien That's a scary reviewer. I don't know much about peer review, but that's like a scary viewer. It's like, oh no, more data analysis, but I have to finish my PhD.

James Zou
Well, well, this here is actually not like providing critiques to the original review, original authors, but more to say here's some new. No, your data is super valuable, and there's actually some new biology that we can find in analysing this data set.

Vivien So others can can could find it too, like say an aeing researcher who's like, I'm really interested in I don't know these astrocytes, and and they're more trying to understand mechanism of what might be behind this, right?

James Zou
That's exactly right. Yeah, because I think the datasets that we often generate now, right, most the modern high throughput techniques are so rich that our initial human analysis only captures a part of the data, right? But there are actually, you know, if you think about it, there are actually a lot of other ideas on how to analyze the data, have time to do like a fraction of the analysis. But the AI agents, because they don't get tired, so they could actually do a lot more in depth, more thorough analysis, and that's actually very useful for the human analyst. Like you know, basically trying to come up with new insights from the data that we have already collected.

Vivien
AI agents can help with new insights from the data that have been collected. As we talked about this, I mentioned another application and paper from his lab called Paper to Agent, which converts research papers into AI agents directly.

The idea is that AI agents can then interact with one another over a paper. Basically, multiple AI agents analyze the paper, then build a model context protocol server, and these MCPs can be hooked up to chatbots. AI agents work through the analysis in a paper to discuss the results. All of which is my simplification and anthropomorphizing, but a link to the paper and the code is in the transcript for you to look at if you like. What's important to Dr. Zou is having a back and forth about science that isn't only about agreeing. As you may have noticed with chatbots, there's a lot of sycophancy.

Scientists don't disagree for the sake of disagreeing, but it's important sometimes to have debate over things in a paper. The concept itself of what a paper is is these days evolving, and these AI agents also bring on new ideas about what a paper might become.

But I do think that you know it's really important to kind of evolve what a paper is and what it can be. A lot of time, the scientists in interviews will say, "Well, we didn't put that in the paper, but it's in supplemental figure, you know, page 76 in this table. And then they have to walk me through that because I'm like, "I don't even know what you're talking about where it is. But so there are things to find in these papers, and also, but I guess sometimes there are kind of incongruencies, and you talked about this, and the Whitehead talked to about the I guess the debate that the agents have, right? You say that they're actually nice, and I mean there is a kind of too much liking and sycophancy in these in the LLM world, but I guess what you're kind of saying is you want this back and forth of, but I think this, but well, I, but anyway, the bot thinks this, yeah, immunology agent says that. You kind of think that that's important for science, right?

James Zou
I think that's critical for science, and because these kinds of debates, that's what really elicits much more robust and more rigorous and creative reasonings, right? Like when one agent has to convince other agents, right? And then I think that leads to better science. And I think a good application of this is that now that we're creating all these paper agents, and that some of the underlying papers are also maybe in conflict with other papers. It's often disagreements in literature.

Vivien
Oh my gosh! And also the ones that aren't published, or that are like, or or just kind of shades of things, and very passionate discussions kind of unfold, or even matters arising, as it's called in exactly.

James Zou
Oh, yeah. So, so one thing that we've tried related to that would be we actually have since we now convert these papers into interactive agents, we can just have two the two agents of like say two papers that disagree with each other just have a debate, right? And then try to convince each other, and then we have like other agents who judge those debate and see who won the debate, and then basically you know which scientific result ends up being more convincing, right? And actually, I looked over some of those debates myself, and I think I found those tips to be really helpful in trying to crystallise what are some of the conflicts, you know.

Vivien
Yeah, I mean that's super important. Also, one thing as I talk to people, maybe in the transcriptomic space, but also elsewhere, this idea that the tools you choose are, you know, the ones that you use because your PI used it, or you just have always used this one and you really like it, or you just like how they keep updating the tools, but the tools sometimes deliver differing results, and then you have data of different types that also don't agree. So this whole kind of idea of harmonising is sometimes hard, and maybe I don't know. People tell me they have a hard time doing this integration.

James Zou
Yeah, I think that's also one of the really good use cases of agents. Right, so now there's there's thousands of tools and more coming out on a weekly basis for for all sorts of let's say genomic or transcriptomic analysis, and it is not possible for individual human researchers to really keep track of all these different tools and figure out which ones to use. But you know that's also where the AI co-sight is agents can help us to manage all the tools, right? And then can, for a given analysis, can suggest which tools are maybe the most appropriate, and then can even just do some initial analysis using those tools for us to evaluate the results,

Vivien
And can it also, for example, I mean, I again, you know, based on my interviews, pipeline breaks because the tool is being updated or something about the update, or there's some missing a library that has been changed. I mean, I guess an agent will come across that too and say, "Oh, this pipeline that I chose doesn't work, and I need to switch to the older version of this tool. Is that kind of something that happens?

James Zou
Yes. Yeah. So we actually did that. Was the Paper2Agent , which under the hood is basically, you know, a team of agents that try to convert papers right into each paper into a interactive agent itself, and in the process of doing that, sometimes it will encounter, like, say, code bases or tools that are maybe are out of date or have inconsistencies in it, and they actually do a pretty good job. The agents do a pretty good job in repairing some of these code bases to make it up to date.
https://arxiv.org/abs/2509.06917

Vivien
Wow, that's huge. I mean, that's huge, particularly in these complex pipelines. That would be amazing. I mean, it also, I guess, the postdoc has left who developed the tool and doesn't, you know, isn't able to keep updating it. And this isn't a professional software environment, and the documentation is all out of whack. Okay, okay, cool. Well, lovely. And I don't want to keep you long. I wanted to ask this other. I guess it's kind of another aspect that you brought up, which is this: when the agent doesn't know. So there's the contrarian one, which I think is super cool that you have. But also, if the information just isn't around, it's not in any databases, it's not at EBI or NIH. Can't you know? Would an agent then? Because you mentioned this. Say, I don't know yet, or I don't know, or the scientists don't know, or let me find out, or what? What does it do in those moments where it says, Hmm, I don't know?

James Zou
Yeah, I mean, I think this is a it's a great question. This is where there's still some limitations in that the current language models might still try to come up with some things, right? It's it's nature like tendencies to not to say I don't know, but there are some ways that we can do to truly make sure that these models are very faithful to the literature, right? So that's if they don't really find anything in the literature, that they they'll have to tell us that they don't they didn't find anything in the literature.

Vivien I mean, it might be in on somebody's hard drive or in somebody's mind, obviously, since a lot of things are happening in parallel. But super cool, I guess. One other question, then I'll let you go to your real day, which is not talking to reporters. So, in theory you can personify or personalise a co-scientist. So not that I wish this for you, but let's say you know your lab could become 500 people because there is an agent James Zhu who is you know, I guess taught with your talks and things that you've written and things that you've said and all of these things. You know, how would you do? You think that that is possible? Because obviously, for mentoring, this would be cool. But also, you know, you don't have to deal with 500 PhD students tugging at you.

James Zou
Yeah, I think that's very interesting. One related experiment we did recently is that we actually created personas based on all these famous scientists, including some historic scientists like Einstein and Feynman and Alan Turing. Right, so I would do this by, you know, for example, like teaching these agents based on the publications and the writing styles and the knowledgebases of these historic scientists. So we basically bring back to life, right? No, sort of like your dream team of allstars, like all the greatest scientists in the world, talent

Vivien Alan Turing, wow, or anybody that you know your your favorite hero, Newton, or yeah,

James Zou
We had Richard Feynman, and then we basically also put them into like you know we personified all these agents and we put them into like kind of a chat forum where they can collaborate and compete on solving open scientific research problems. What's really quite amazing is that within about 30 minutes, they actually took one of these famous Erdö s math problems and actually discovered the best, currently the best new solution to that problem.

Vivien Oh, I think I saw that on GitHub. You've got software, didn't I? The optimization. I don't know enough about math to understand the deep. Yeah,

James Zou
This is like the Erdös. It's called the Erdös minimum overlap problem, which is like a f a mous problem that a lot of people have and researchers have spent time on, and we we re actually quite amazed that you know this team of these historic scientists agents that are autonomously they they work together and compete and share notes, and then within less than an hour, they're able to actually come up with a better solution than that anyone has thought before.

Vivien
That is so cool. That is really really cool. I mean, it's it's really bringing history to life. You haven't said if you would mind the James Z AI agent, but I guess what you're saying is that people can build it, can build their mentor or their hero or their colleague or use things possibly from Slack chats and Reddits and all kinds of things.

James Zou
I think so. I think that's definitely will become more and more common. What I would say is that what I really like about being a professor is still the the in person, like the person to person mentorship and interactions, of course, my students, right? So that's physical interactions, right? I would still definitely want to maintain and keep, but I do see that there's some the other ways that let's say a James Zou agent could be useful. You know, maybe like in let's say, if I have research projects, I don't have time to do myself, right? Then I'd love to have my agent help me to start to work on some of those projects.

Vivien
Fabulous! Ah, this has been great fun. Thank you so much, James.

James
Well, well, thank you. I really enjoyed our conversations as always.

Vivien
Take care. Bye.

James Zou
Take care. Have a good day. Bye.

Vivien
And that was Conversations with scientists. Today's episode was with Dr. James Zou from Stanford University. The music in this podcast is Honey Mustard by Randy Sharpe, licensed from Artlist.io. And I just wanted to say, because there's confusion about these things sometimes, Dr. Zou didn't pay to be in this podcast, nor did Stanford University pay for it to be produced. This is independent journalism that I produce in my living room. I'm Vivien Marx. Thanks for listening.