专业介绍
Data Science is often viewed as the confluence of (1) Computer and Information Sciences (2) Statistical Sciences, and (3) Domain Expertise. These three pillars are not symmetric: the first two together represent the core methodologies and the techniques used in Data Science, while the third pillar is the application domain to which this methodology is applied. In this program, core data science training is focused on the first two pillars, along with practice in applying their skills to address problems in application domains.
We characterize the required Data Science skills in two categories: statistical skills, such as those taught by the Statistics and Biostatistics departments, and computational skills, such as those taught by the Computer Science and Engineering Division and the School of Information. The design of the program is to require every student to receive balanced training in both areas. To create an academic plan that achieves this balance, and to foster a greater sense of shared community, we do not intend to offer any sub-plans or tracks within the proposed degree program. Rather, we will expect graduates of this program to understand data representation and analysis at an advanced level.
With the MS in Data Science all students will be able to:
- identify relevant datasets
- apply the appropriate statistical and computational tools to the dataset to answer questions posed by individuals, organizations or governmental agencies
- design and evaluate analytical procedures appropriate to the data
- implement these efficiently over large heterogeneous data sets in a multi-computer environment