LamaLab · Benchmark
ChemBench
Are large language models superhuman chemists?
Topic columns
Time axis
Show
Each dot is a model at its release date; the stepped line is the best overall
score achieved up to that point (state of the art). Hover a dot for details.
Best score per topic over time. Click any topic to expand.
Each model family's score over time. A red ▼ marks a release that scored lower than the family's previous one. Click any topic below to expand.
Overall score per model family over time. Lines connect a family's successive releases; a red ▼ marks a release that scored lower than the family's previous one.