Where disciplines collaborate and research meets education: tackling real-world problems with data, tools, methods — and preparing the next generation of polymath researchers.

A large language model is a statistical machine, which makes statistical literacy a basic requirement for everyone who uses one.
Writing is thinking. Coding is thinking. Exploratory data analysis is thinking, and so is creating a good visualization. Working with AI can be thinking too — when you actively engage.
The aiX Faculty Fellowship Program returns for 2026–2027. This year’s fellowship continues to support faculty in designing curricular innovations on teaching with AI and will focus more on curriculum development on teaching about AI within the context of their own discipline. The call for applications is forthcoming — we aim to open applications in early October 2026.
I am teaching Introduction to Statistical Reasoning this fall, and while preparing for the course I found myself returning to a basic question: what is statistics really about?

The AI for Social Good and Society (AI4SGS) Initiative is a bold interdisciplinary effort to apply artificial intelligence to some of the world’s most pressing social and public health challenges.

We explore genAI tools to lower the barriers in Climate Data Science in collaboration with AWS.
Currect subprojects include: Knowledge graph construction and expansion. Development of agents for data acquision, analysis, modeling, visualization, etc. Evaluation through case studies.

Learning the Earth with Artificial Intelligence and Physics (LEAP) is an NSF Science and Technology Center (STC) launched in 2021. LEAP’s mission is to increase the reliability, utility, and reach of climate projections through the integration of climate and data science.

A long-time collaboration between Professors Tian Zheng and Professor Maria Uriarte on using machine learning to unlock potentials of new data types to understand the impact of climate change on tropical forests.

The Collaboratory is both a set of “data science in context” educational approaches, as well as a meta-model for an accelerator program that allows different institutions to respond flexibly to their own disciplinary heterogeneity in terms of data science educational needs. The novelty of the Collaboratory lies in its crowd-sourcing approach to creating new data science pedagogy and its ability to kindle transdisciplinary collaboration in doing so. Read our Havard Data Science Review article to learn more.

Applied Data Science at Columbia is a project-based learning course that started in 2016. It employs the common task framework and runs 5 mini project cycles during one semester to give students a broad exposure to various areas in data science. Projects are developed and updated each year, drawing inspirations from active research, challenges and interesting public datasets.

A core and active research area of the TZstats lab.