FONDAMENTI DI ANALISI DATI E LABORATORIO
Module LABORATORIO

Academic Year 2026/2027 - Teacher: ANTONINO FURNARI

Expected Learning Outcomes

  1. Knowledge and understanding: Students will acquire knowledge of the software tool ecosystem for data analysis in Python and understand how the techniques presented in the theoretical module translate into concrete operations on real datasets, including the parameters they require and the form of their output.
  2. Applying knowledge and understanding: Students will acquire technical skills for building, managing, and analyzing real datasets. They will be able to load and clean a dataset, explore it, produce appropriate visualizations, apply the statistical techniques and predictive models studied using reference libraries, and evaluate the results obtained—with the goal of building models and decision-support systems.
  3. Making judgements: Students will be able to independently organize an analysis workflow on an unseen dataset, choosing the appropriate tools for each phase, critically interpreting the output generated by libraries, and identifying anomalous or unreliable results.
  4. Communication skills: Students will be able to draft comprehensive, visually appropriate reports to communicate data analysis and exploration results correctly and effectively, as well as orally present their work while justifying the choices made.
  5. Learning skills: Students will develop the necessary skills to independently stay up to date on techniques, software, and programming languages for data analysis, learning to consult library documentation and evaluate new tools to ensure continuous learning beyond the course.

Course Structure

Practical hands-on lab sessions held in the classroom, where the techniques studied are demonstrated and applied through code examples and guided analyses on real datasets. Sessions in this module are tightly intertwined with the theoretical lectures of the Fundamentals of Data Analysis module: each topic is addressed from an implementation standpoint on the same day or shortly thereafter. The two components form a unified learning path and are assessed through a single exam.

Work is conducted in Python using notebooks. The code presented in class is made available to students, who are encouraged to re-run, modify, and apply it to different datasets.

Should the course be delivered in a hybrid or remote format, necessary variations to the plan outlined above may be introduced to fulfill the syllabus requirements.

Required Prerequisites

Basic programming skills, mathematical analysis, and linear algebra are required. Prior knowledge of the Python programming language or the libraries used is not required, as these tools are introduced during the course.

Attendance of Lessons

Attending lectures is not mandatory, but strongly recommended.

Detailed Course Content

The module follows the structure of the theoretical counterpart, divided into two parts (Data Analysis; Predictive Techniques and Data Representation), and develops its practical application.

Data Analysis Tools in Python

  • Work environment: Notebooks, managing an analysis project, reproducibility of the analysis.
  • Elements of numerical computing with NumPy: Arrays, vectorized operations, indexing.
  • Tabular data manipulation with pandas: Loading from files and the web, selection and filtering, grouping and aggregation, joining tables.
  • Overview of libraries used in the course and criteria for navigating their documentation.

Data Preparation and Exploration

  • Dataset construction from real sources: Initial inspection and quality checks.
  • Data cleaning: Handling missing values, identifying and treating outliers, correcting types and formats.
  • Data transformation: Standardization and normalization, encoding categorical variables.
  • Calculation of descriptive statistics and correlation measures on real datasets.
  • Data visualization with Matplotlib and Seaborn: Histograms, boxplots, scatter plots, hexbin plots, density maps, contour plots, scatter matrices, regression plots; customizing charts and selecting representations based on the message.

Data Analysis: Probability, Inference, and Regression Models

  • Probability distributions with scipy.stats: Sampling, density, cumulative distribution functions; simulation to illustrate sampling variability and the central limit theorem.
  • Calculating confidence intervals and performing statistical tests; reading and interpreting library output; calculating effect size measures.
  • Implementing bootstrap and permutation tests.
  • Estimating regression models with statsmodels: Reading result tables, coefficients, standard errors, significance, goodness of fit, residual diagnostics.
  • Multiple regression in practice: Effect of adding control variables on estimated coefficients, interaction terms and polynomial terms, encoding categorical variables in model formulas.
  • Logistic regression: Estimation and interpretation of coefficients in terms of odds.

Predictive Techniques with scikit-learn

  • Structure of the scikit-learn library: Estimators, transformers, pipelines; architectural differences compared to statsmodels.
  • Data splitting: Training, validation, and test sets; cross-validation; hyperparameter tuning.
  • Diagnosing overfitting: Comparing performance on training vs. test sets across model complexity levels; using regularization for control.
  • Predictive models: Linear and logistic regression evaluated on unseen data, K-NN, generative classifiers (QDA, LDA, Naive Bayes); visualizing decision boundaries.
  • Calculation and interpretation of evaluation metrics for regression and classification; model comparison.
  • Scores, thresholds, and curves: Constructing ROC and precision-recall curves from model scores, impact of threshold shifting on error types, threshold selection based on costs, comparing both curves on imbalanced datasets.

Data Representation

  • Clustering with K-Means: Criteria for choosing the number of clusters and visual results inspection.
  • Density estimation and Gaussian Mixture Models (GMM).
  • PCA: Computation, interpretation of components, application in visualization and as a step in a pipeline.

Project and Results Communication

  • Organizing an end-to-end analysis workflow on a medium-to-large dataset: From the initial question to the final conclusion.
  • Structuring a data analysis report: Selecting and refining visualizations designed for communication.
  • Preparing and delivering the presentation of results.

Textbook Information

The primary reference for the module consists of the notebooks and lab materials provided by the instructor, accessible via Microsoft Teams (Team code: i87g4nb) and through the website http://antoninofurnari.github.io/fadlecturenotes/.

Recommended texts for further reading:

  1. James, G., Witten, D., Hastie, T., Tibshirani, R. An Introduction to Statistical Learning, with Applications in Python, 2023. https://www.statlearning.com
  2. Knaflic, C. N. Storytelling with Data, John Wiley & Sons, 2025.
  3. Official documentation for the libraries used: numpy, pandas, matplotlib, seaborn, scipy, statsmodels, scikit-learn.

For the theoretical foundations of the topics covered, please refer to the texts listed in the syllabus for the Fundamentals of Data Analysis module.

Course Planning

 SubjectsText References
1Work environment, notebooks, and introduction to numpyLecture notes; [3]
2Tabular data manipulation with pandasLecture notes; [3]
3Data cleaning and preparationLecture notes; [3]
4Descriptive statistics and correlation measures on real datasetsLecture notes;
5Data visualization with matplotlib and seabornLecture notes; [2]
6Probability distributions and simulation with scipy.statsLecture notes; [3]
7Confidence intervals, statistical tests, bootstrap, and permutation testsLecture notes; [1]
8Regression with statsmodels: estimation, reading output, significance, diagnosticsLecture notes; [1]
9Multiple regression in practice: control variables, interactions, polynomial terms. Logistic regression and oddsLecture notes; [1]
10Introduction to scikit-learn: pipelines, cross-validation, overfitting diagnosis, and regularizationLecture notes; [1]
11Predictive models: linear and logistic regression for prediction, K-NN, QDA, LDA, Naive Bayes; evaluation metricsLecture notes; [1]
12Clustering, density estimation, and PCALecture notes; [1]

Learning Assessment

Learning Assessment Procedures

The exam is divided into the following tests:

  1. A written exam designed to assess the student's theoretical understanding of the topics covered in the course, from both a theoretical and methodological perspective. The exam is graded on a scale of thirty.
  2. A project assigned by the instructor and carried out independently by the student, aimed at evaluating practical skills in data analysis and communication of results. The project is presented to the instructor through a presentation and graded on a scale of thirty.

Students with disabilities and/or DSA must contact the teacher, the CInAP representative of the DMI (Prof. Daniele) and CInAP well in advance of the exam date to communicate that they intend to take the exam using the appropriate compensatory measures.

Two written in itinere exams are scheduled during the course. Passing both tests grants exemption from the final written exam.

The final grade is obtained by means of a weighted average between the marks obtained in the two tests with weights of 40% for the written test and 60% for the project.

The assessment of learning can also be conducted remotely if the conditions require it.

The grading of each test is expressed on a scale of thirty points according to the following scheme:

Score 29-30 with honors

The student has a deep understanding of the concepts and techniques of data analysis. They can promptly analyze data analysis problems, identifying the most suitable data analysis techniques for the given problem independently and critically, and indicating the most suitable methodological practices for their application. They have excellent communication skills and language proficiency.

Score 26-28

The student has a good understanding of the concepts and techniques of data analysis. They can analyze data analysis problems, identifying appropriate data analysis techniques for the given problem and indicating suitable methodological practices for their application. They have good communication skills and language proficiency.

Score 22-25

The student has a fair knowledge of the concepts and techniques of data analysis, although it may be limited to the main topics. They can analyze data analysis problems, albeit not always in a linear manner, identifying suitable data analysis techniques for the given problem. They have fair communication skills and language proficiency.

Score 18-21

The student has minimal knowledge of the concepts and techniques of data analysis. They have limited ability to analyze data analysis problems. They have sufficient communication skills, although not always appropriate language proficiency.

Examination not passed

The student does not possess the minimum required knowledge of the main content of the course. Their ability to use specific language is very poor or nonexistent, and they are unable to independently apply the acquired knowledge.

Examples of frequently asked questions and / or exercises

The data analysis project is generally based on medium-to-large datasets available online. The project typically requires students to:

  • Load a dataset from a real-world source, evaluate its quality, and prepare it for analysis, justifying the choices made regarding missing data and outliers;
  • Conduct an exploratory analysis and produce a set of visualizations that describe the data structure and the relationships between variables;
  • Formulate one or more questions about the data and answer them using appropriate inferential tools, reporting the associated uncertainties;
  • Build a predictive model tailored to the problem, evaluate its performance using appropriate metrics, and compare it against at least one alternative;
  • Write a report and deliver an oral presentation of the work, justifying the methodological choices and discussing the limitations of the analysis.