Scientists' First Exam released

Jun 12, 2025 · 1 min read
blog

We released Scientists’ First Exam, a large-scale multimodal scientific benchmark for scientific discovery scenarios. The dataset is available on Hugging Face.

Wenlong Zhang
Authors
Young Researcher
Agent Center

I am a Young Researcher of Shanghai AI Laboratory, working with Prof. Bo Zhang, Prof. Wanli Ouyang and Prof. Lei Bai. Before that, I got the PhD degree from Hong Kong Polytechnic University, working with Prof. Xiao-Ming Wu. I also interned at XPixel group in Shanghai AI Laboratory and SIAT-CAS, working with Prof. Chao Dong and Prof. Yu Qiao. In 2018, I got the Master degree from the Beijing Institute of Technology, supervised by Prof. Weidong Hu.

Currently, I lead a team engaged in Agent model Alignment, Post-training and Evaluation. Recently, my primary areas of focus include:

  • Long-horizon agent evaluation: How do we formulate the workflow of scientific and real-world scenarios with agent models into a unified benchmark?
  • Agent model alignment: How can we align agent models with human preferences, scientific objectives, and domain constraints while ensuring reliability, generalization, and controllability?
  • Agent model post-training: How can we develop scalable data and optimization strategies to improve the planning, reasoning, tool use, and long-horizon decision-making capabilities of agent models?