OpenAI has introduced LifeSciBench, a new benchmark built for life-science work rather than just general biology questions. The project lays out 750 expert-authored tasks drawn from seven life-science workflows and seven biological domains, according to OpenAI.
OpenAI says the benchmark is grounded in the judgment of practicing life scientists, specifically people with Ph.D.-level training and biotech or pharma experience. That matters because the goal is to test whether AI can handle the kinds of steps and decisions that show up in real research settings—where evidence, constraints, and workflow details often matter as much as the final answer.
More than “can it explain biology?”
In OpenAI’s description, LifeSciBench is meant to measure whether AI systems can support realistic life science research tasks. The emphasis is on task support within workflows, not simply producing textbook-style responses about biological topics.
Practically speaking, that framing suggests the benchmark is trying to evaluate whether an AI can do useful work across different parts of a life-science process—something closer to how labs and teams actually operate. OpenAI does not present claims here about how any specific system performs, but the benchmark structure itself could influence what researchers choose to build or test next.
As more AI tools move beyond chat and into domain-specific help, benchmarks like this are often the gatekeepers. They can shape what “good performance” looks like for life-science teams, and they can also highlight where systems still fall short when the job is more complex than explaining a concept.
Source: OpenAI

