Project: ArXivLab: A Platform for Developing and Evaluating Exploratory Tools for the Scientific Literature

January 11, 2019 in Project 2019

Through ArXivLab we aim to develop the next generation recommender systems for the scientific literature using statistical machine learning approaches. In collaboration with ArXiv we are currently developing a new scholarly literature browser which will be able to extract knowledge implicit in the mathematical and scientific literature, offer advanced mathematical search capabilities and provide personalized recommendations.

Project: Measuring Liberal Arts: Creating an Index for Higher Education

January 10, 2019 in Project 2019

This project works with a novel corpus of text-based school data to develop a multi-dimensional measure of the degree to which American colleges and universities offer a liberal arts education. We seek a data scientist for various tasks on a project that uses analysis of multiple text corpora to better understand the liberal arts. This is an ongoing three-year project with opportunities for future collaborations, academic publications, and developing and improving existing data science and machine learning skills.

Project: Project ToxicDocs

January 23, 2018 in Project 2018

We house the world’s largest dataset of once-secret documents on industrial pollution, unleashed from the vaults of corporations like DuPont, Dow, and Monsanto in toxic tort litigation. We are applying data science methods to analyze and render this material useable to a broad audience.

Project: Data Science and the regulation of financial markets (application closed)

January 22, 2018 in Project 2018

The development of computational data science techniques in natural language processing (NLP) and machine learning (ML) algorithms to analyze large and complex textual information opens new avenues to study intricate processes, such as government regulation of financial markets, at a scale unimaginable even a few years ago. This project develops scalable NLP and ML algorithms (classification, clustering and ranking methods) that automatically classify laws into various codes/labels, rank feature sets based on use case, and induce best structured representation of sentences for various types of computational analysis.

Project: ArXivLab: A Platform for Developing and Evaluating Exploratory Tools for the Scientific Literature

Project: Measuring Liberal Arts: Creating an Index for Higher Education

Project: Project ToxicDocs

Project: Data Science and the regulation of financial markets (application closed)

Columbia Data Science Institute (DSI) Scholars Program