FONDAMENTI DI ANALISI DATI E LABORATORIOModule LABORATORIO
Academic Year 2026/2027 - Teacher: ANTONINO FURNARIExpected Learning Outcomes
- Knowledge and understanding: Students will acquire knowledge of the software tool ecosystem for data analysis in Python and understand how the techniques presented in the theoretical module translate into concrete operations on real datasets, including the parameters they require and the form of their output.
- Applying knowledge and understanding: Students will acquire technical skills for building, managing, and analyzing real datasets. They will be able to load and clean a dataset, explore it, produce appropriate visualizations, apply the statistical techniques and predictive models studied using reference libraries, and evaluate the results obtained—with the goal of building models and decision-support systems.
- Making judgements: Students will be able to independently organize an analysis workflow on an unseen dataset, choosing the appropriate tools for each phase, critically interpreting the output generated by libraries, and identifying anomalous or unreliable results.
- Communication skills: Students will be able to draft comprehensive, visually appropriate reports to communicate data analysis and exploration results correctly and effectively, as well as orally present their work while justifying the choices made.
- Learning skills: Students will develop the necessary skills to independently stay up to date on techniques, software, and programming languages for data analysis, learning to consult library documentation and evaluate new tools to ensure continuous learning beyond the course.
Course Structure
Practical hands-on lab sessions held in the classroom, where the techniques studied are demonstrated and applied through code examples and guided analyses on real datasets. Sessions in this module are tightly intertwined with the theoretical lectures of the Fundamentals of Data Analysis module: each topic is addressed from an implementation standpoint on the same day or shortly thereafter. The two components form a unified learning path and are assessed through a single exam.
Work is conducted in Python using notebooks. The code presented in class is made available to students, who are encouraged to re-run, modify, and apply it to different datasets.
Should the course be delivered in a hybrid or remote format, necessary variations to the plan outlined above may be introduced to fulfill the syllabus requirements.
Required Prerequisites
Attendance of Lessons
Attending lectures is not mandatory, but strongly recommended.
Detailed Course Content
The module follows the structure of the theoretical counterpart, divided into two parts (Data Analysis; Predictive Techniques and Data Representation), and develops its practical application.
- Work environment: Notebooks, managing an analysis project, reproducibility of the analysis.
- Elements of numerical computing with NumPy: Arrays, vectorized operations, indexing.
- Tabular data manipulation with pandas: Loading from files and the web, selection and filtering, grouping and aggregation, joining tables.
- Overview of libraries used in the course and criteria for navigating their documentation.
- Dataset construction from real sources: Initial inspection and quality checks.
- Data cleaning: Handling missing values, identifying and treating outliers, correcting types and formats.
- Data transformation: Standardization and normalization, encoding categorical variables.
- Calculation of descriptive statistics and correlation measures on real datasets.
- Data visualization with Matplotlib and Seaborn: Histograms, boxplots, scatter plots, hexbin plots, density maps, contour plots, scatter matrices, regression plots; customizing charts and selecting representations based on the message.
- Probability distributions with scipy.stats: Sampling, density, cumulative distribution functions; simulation to illustrate sampling variability and the central limit theorem.
- Calculating confidence intervals and performing statistical tests; reading and interpreting library output; calculating effect size measures.
- Implementing bootstrap and permutation tests.
- Estimating regression models with statsmodels: Reading result tables, coefficients, standard errors, significance, goodness of fit, residual diagnostics.
- Multiple regression in practice: Effect of adding control variables on estimated coefficients, interaction terms and polynomial terms, encoding categorical variables in model formulas.
- Logistic regression: Estimation and interpretation of coefficients in terms of odds.
- Structure of the scikit-learn library: Estimators, transformers, pipelines; architectural differences compared to statsmodels.
- Data splitting: Training, validation, and test sets; cross-validation; hyperparameter tuning.
- Diagnosing overfitting: Comparing performance on training vs. test sets across model complexity levels; using regularization for control.
- Predictive models: Linear and logistic regression evaluated on unseen data, K-NN, generative classifiers (QDA, LDA, Naive Bayes); visualizing decision boundaries.
- Calculation and interpretation of evaluation metrics for regression and classification; model comparison.
- Scores, thresholds, and curves: Constructing ROC and precision-recall curves from model scores, impact of threshold shifting on error types, threshold selection based on costs, comparing both curves on imbalanced datasets.
- Clustering with K-Means: Criteria for choosing the number of clusters and visual results inspection.
- Density estimation and Gaussian Mixture Models (GMM).
- PCA: Computation, interpretation of components, application in visualization and as a step in a pipeline.
- Organizing an end-to-end analysis workflow on a medium-to-large dataset: From the initial question to the final conclusion.
- Structuring a data analysis report: Selecting and refining visualizations designed for communication.
- Preparing and delivering the presentation of results.
Textbook Information
The primary reference for the module consists of the notebooks and lab materials provided by the instructor, accessible via Microsoft Teams (Team code: i87g4nb) and through the website http://antoninofurnari.github.io/fadlecturenotes/.
- James, G., Witten, D., Hastie, T., Tibshirani, R. An Introduction to Statistical Learning, with Applications in Python, 2023. https://www.statlearning.com
- Knaflic, C. N. Storytelling with Data, John Wiley & Sons, 2025.
- Official documentation for the libraries used: numpy, pandas, matplotlib, seaborn, scipy, statsmodels, scikit-learn.
Course Planning
| Subjects | Text References | |
|---|---|---|
| 1 | Work environment, notebooks, and introduction to numpy | Lecture notes; [3] |
| 2 | Tabular data manipulation with pandas | Lecture notes; [3] |
| 3 | Data cleaning and preparation | Lecture notes; [3] |
| 4 | Descriptive statistics and correlation measures on real datasets | Lecture notes; |
| 5 | Data visualization with matplotlib and seaborn | Lecture notes; [2] |
| 6 | Probability distributions and simulation with scipy.stats | Lecture notes; [3] |
| 7 | Confidence intervals, statistical tests, bootstrap, and permutation tests | Lecture notes; [1] |
| 8 | Regression with statsmodels: estimation, reading output, significance, diagnostics | Lecture notes; [1] |
| 9 | Multiple regression in practice: control variables, interactions, polynomial terms. Logistic regression and odds | Lecture notes; [1] |
| 10 | Introduction to scikit-learn: pipelines, cross-validation, overfitting diagnosis, and regularization | Lecture notes; [1] |
| 11 | Predictive models: linear and logistic regression for prediction, K-NN, QDA, LDA, Naive Bayes; evaluation metrics | Lecture notes; [1] |
| 12 | Clustering, density estimation, and PCA | Lecture notes; [1] |
Learning Assessment
Learning Assessment Procedures
- A written exam designed to assess the student's theoretical understanding of the topics covered in the course, from both a theoretical and methodological perspective. The exam is graded on a scale of thirty.
- A project assigned by the instructor and carried out independently by the student, aimed at evaluating practical skills in data analysis and communication of results. The project is presented to the instructor through a presentation and graded on a scale of thirty.
Examples of frequently asked questions and / or exercises
The data analysis project is generally based on medium-to-large datasets available online. The project typically requires students to:
- Load a dataset from a real-world source, evaluate its quality, and prepare it for analysis, justifying the choices made regarding missing data and outliers;
- Conduct an exploratory analysis and produce a set of visualizations that describe the data structure and the relationships between variables;
- Formulate one or more questions about the data and answer them using appropriate inferential tools, reporting the associated uncertainties;
- Build a predictive model tailored to the problem, evaluate its performance using appropriate metrics, and compare it against at least one alternative;
- Write a report and deliver an oral presentation of the work, justifying the methodological choices and discussing the limitations of the analysis.