FONDAMENTI DI ANALISI DATI E LABORATORIOModule FONDAMENTI DI ANALISI DATI
Academic Year 2026/2027 - Teacher: ANTONINO FURNARIExpected Learning Outcomes
- Knowledge and Understanding: The student will gain a solid understanding of the fundamental principles needed to collect, organize, model, analyze, and interpret data. This will be achieved through the presentation of a theoretical-mathematical framework and numerous examples of its application to real datasets. The student will develop a deep understanding of the conceptual foundations of data analysis.
- Applying Knowledge and Understanding: The student will acquire technical skills for constructing, managing, and analyzing real datasets, with the goal of building models and decision support systems. They will be able to apply the acquired knowledge to solve practical problems using tools and techniques for data analysis.
- Making Judgements: The student will be able to independently choose the most appropriate techniques for solving a data analysis problem, evaluating their pros and cons. They will be capable of justifying their choices and critically assessing various methodologies for data analysis and knowledge extraction.
- Communication Skills: The student will be trained to produce complete, rigorous, and visually appropriate reports that effectively and correctly communicate the results of data analysis and exploration. Conclusions will be clearly justified and communicated effectively to both technical and non-technical audiences.
- Learning Skills: The student will develop the necessary skills to update themselves independently on the use of techniques, software, and programming languages useful for data analysis, ensuring continuous learning even after the course ends.
Course Structure
In-person classroom lectures dedicated to the theoretical and methodological coverage of the course topics. The lectures in this module are tightly intertwined with those of the Laboratory module: topics covered from a theoretical perspective are revisited and applied—either on the same day or shortly thereafter—through code examples and guided analyses on real-world datasets. The two components form a unified learning path and are evaluated through a single exam.
Should the course be delivered in a blended or remote format, necessary adjustments to the aforementioned structure may be introduced in order to complete the syllabus as planned.
Required Prerequisites
Attendance of Lessons
Attending lectures is not mandatory, but strongly recommended.
Detailed Course Content
- Data Analysis
- Predictive Techniques and Data Representation
- Overview of data analysis: Main types, purposes and applications, data analysis examples
- Different data types: Nominal, ordinal, interval, and ratio data
- Data collection techniques: Surveys, experiments, observational studies, sampling
- Difference between sample and population
- Data pre-processing techniques: Data cleaning, handling missing data, data standardization, categorical variable encoding (dummy variables), noise reduction in data (filtering, outlier removal, normalization)
- Using probability for data analysis: Basic concepts of probability (joint, marginal, conditional probability, independence and conditional independence), Bayes' theorem and its application in data analysis, discrete, continuous, cumulative probability distributions. Notable probability distributions
- Measures of central tendency (mean, median, and mode), measures of dispersion (variance, standard deviation, quartiles, and interquartile range)
- Covariance, correlation measures between variables
- Data visualization techniques: Pie charts, histograms, boxplots, scatterplots, hexbin plots, density maps, contour plots, scatter matrices, regression plots
- Use of inferential data analysis tools: Confidence intervals, significance levels, and statistical tests. Group comparisons, effect size, multiple testing problem. Computational methods: bootstrap and permutation tests
- Linear regression and logistic regression for studying relationships between variables: Estimation and interpretation of coefficients, standard errors and statistical significance of coefficients, goodness of fit and residual diagnostics, control variables, interaction terms and polynomial terms, variable selection and backward elimination, interpretation of logistic regression coefficients in terms of odds, multinomial regression
- Correlation and causality: Difference between association and causal relationship, randomized controlled trials and observational studies, experimental design, confounding variables and their control via regression, limitations of causal conclusions drawn from observational data
- Fundamental concepts of predictive analysis: From describing observed data to predicting unobserved data. Training, validation, and test sets; cross-validation. Parameters and hyperparameters. Parametric and non-parametric methods. Linear and non-linear models
- Generalization capability: Overfitting and underfitting as forms of model fitting to sample peculiarities rather than population regularities, bias and variance and their connection to model complexity, controlling complexity via variable selection and regularization
- Linear regression as a predictive method: Evaluation metrics for regression problems (mean squared error and mean absolute error), evaluation on unobserved data, difference between a model that describes observed data well and one that predicts new data well
- Classification techniques. Evaluating the performance of a classification model: Confusion matrix, precision, recall, and F1 score; behavior of metrics in the presence of imbalanced classes. K-Nearest Neighbors (KNN). Logistic and multinomial regression as predictive methods
- The classifier as a decision rule on a score: Distinction between the quality of evidence provided by the data and the choice of the operational threshold. ROC curves and area under the curve (AUC) for studying the trade-off between the two types of errors as the threshold varies, interpretation of AUC in terms of observation ranking, threshold selection based on error costs, comparison between ROC curves and precision-recall curves in the presence of imbalanced classes
- Generative classifiers: MAP decision rule as an application of Bayes' theorem, class-conditional density estimation, quadratic discriminant analysis (QDA) and linear discriminant analysis (LDA), Naive Bayes. Effect of different assumptions on the decision boundary and on the number of parameters to estimate
- Features, representation functions, feature spaces, metrics
- Clustering techniques: Definitions and K-Means
- Fitting Gaussians to data, Maximum Likelihood, Gaussian Mixture Models
- Non-parametric density estimation using Kernel Density Estimation
- Principal Component Analysis (PCA)
Textbook Information
- Peck, Roxy, Chris Olsen, and Jay L. Devore. Introduction to statistics and data analysis. Cengage Learning, 2015.
- James, Gareth Gareth Michael. An introduction to statistical learning: with applications in Python, 2023.https://www.statlearning.com
- Bishop, Christopher M. "Machine Learning. Machine learning, 2006. https://www.microsoft.com/en-us/research/publication/pattern-recognition-machine-learning/
- Hernán, Miguel A., and James M. Robins. Causal inference, 2010. https://www.hsph.harvard.edu/miguel-hernan/causal-inference-book/
- Knaflic, Cole Nussbaumer. Storytelling with data: A data visualization guide for business professionals. John Wiley & Sons, 2025.
Course Planning
| Subjects | Text References | |
|---|---|---|
| 1 | Introduction to the course | [1,2]; lecture notes |
| 2 | Main Data Analysis Concepts | [1]; lecture notes |
| 3 | Descriptive Statistics and Graphical Representation of data | [1]; lecture notes |
| 4 | Uncertainty and data as the observation of random events | [1,3]; lecture notes |
| 5 | Probability for Data Analysis | [1,3]; lecture notes |
| 6 | Associations of Two Variables | [1]; lecture notes |
| 7 | Use of statistical inference in data analysis | [1]; lecture notes |
| 8 | Fundamentals of Causal Data Analysis | [5]; lecture notes |
| 9 | Storytelling with data | [4]; lecture notes |
| 10 | Predictive Data Analysis | [3]; lecture notes |
| 11 | Probabilistic Models for Classification | [3]; lecture notes |
| 12 | Clustering & Density estimation | [3]; lecture notes |
| 13 | Dimensionality Reduction and Principal Component Analysis | [3]; lecture notes |
Learning Assessment
Learning Assessment Procedures
- A written exam designed to assess the student's theoretical understanding of the topics covered in the course, from both a theoretical and methodological perspective. The exam is graded on a scale of thirty.
- A project assigned by the instructor and carried out independently by the student, aimed at evaluating practical skills in data analysis and communication of results. The project is presented to the instructor through a presentation and graded on a scale of thirty.
Examples of frequently asked questions and / or exercises
The data analysis project is generally based on medium-large datasets obtainable on the internet.
Examples of typical exam questions:
- Define the classification problem, discuss the differences with respect to the regression problem and give practical examples.
- Explain the K-NN algorithm for classification. Discuss the effect of parameter K on algorithm performance. Give graphical examples of how the algorithm works and the effect of K.
- Discuss evaluation measures for classification problems: accuracy, confusion matrix, precision, recall and F1 score. The pros and cons of the measures considered are discussed, also in relation to the characteristics of the test dataset.
- Illustrate the main techniques useful for studying the correlation between variables.