FONDAMENTI DI ANALISI DATI E LABORATORIO
Module FONDAMENTI DI ANALISI DATI

Academic Year 2026/2027 - Teacher: ANTONINO FURNARI

Expected Learning Outcomes

  1. Knowledge and Understanding: The student will gain a solid understanding of the fundamental principles needed to collect, organize, model, analyze, and interpret data. This will be achieved through the presentation of a theoretical-mathematical framework and numerous examples of its application to real datasets. The student will develop a deep understanding of the conceptual foundations of data analysis.
  2. Applying Knowledge and Understanding: The student will acquire technical skills for constructing, managing, and analyzing real datasets, with the goal of building models and decision support systems. They will be able to apply the acquired knowledge to solve practical problems using tools and techniques for data analysis.
  3. Making Judgements: The student will be able to independently choose the most appropriate techniques for solving a data analysis problem, evaluating their pros and cons. They will be capable of justifying their choices and critically assessing various methodologies for data analysis and knowledge extraction.
  4. Communication Skills: The student will be trained to produce complete, rigorous, and visually appropriate reports that effectively and correctly communicate the results of data analysis and exploration. Conclusions will be clearly justified and communicated effectively to both technical and non-technical audiences.
  5. Learning Skills: The student will develop the necessary skills to update themselves independently on the use of techniques, software, and programming languages useful for data analysis, ensuring continuous learning even after the course ends.
-->

Course Structure

In-person classroom lectures dedicated to the theoretical and methodological coverage of the course topics. The lectures in this module are tightly intertwined with those of the Laboratory module: topics covered from a theoretical perspective are revisited and applied—either on the same day or shortly thereafter—through code examples and guided analyses on real-world datasets. The two components form a unified learning path and are evaluated through a single exam.

Should the course be delivered in a blended or remote format, necessary adjustments to the aforementioned structure may be introduced in order to complete the syllabus as planned.

Required Prerequisites

Sono richieste competenze di base di programmazione, analisi matematica e algebra lineare.

Attendance of Lessons

Attending lectures is not mandatory, but strongly recommended.

Detailed Course Content

The course is structured into two main modules:

  • Data Analysis
  • Predictive Techniques and Data Representation

The following paragraphs detail the contents of the various modules.


Data Analysis

  • Overview of data analysis: Main types, purposes and applications, data analysis examples
  • Different data types: Nominal, ordinal, interval, and ratio data
  • Data collection techniques: Surveys, experiments, observational studies, sampling
  • Difference between sample and population
  • Data pre-processing techniques: Data cleaning, handling missing data, data standardization, categorical variable encoding (dummy variables), noise reduction in data (filtering, outlier removal, normalization)
  • Using probability for data analysis: Basic concepts of probability (joint, marginal, conditional probability, independence and conditional independence), Bayes' theorem and its application in data analysis, discrete, continuous, cumulative probability distributions. Notable probability distributions
  • Measures of central tendency (mean, median, and mode), measures of dispersion (variance, standard deviation, quartiles, and interquartile range)
  • Covariance, correlation measures between variables
  • Data visualization techniques: Pie charts, histograms, boxplots, scatterplots, hexbin plots, density maps, contour plots, scatter matrices, regression plots
  • Use of inferential data analysis tools: Confidence intervals, significance levels, and statistical tests. Group comparisons, effect size, multiple testing problem. Computational methods: bootstrap and permutation tests
  • Linear regression and logistic regression for studying relationships between variables: Estimation and interpretation of coefficients, standard errors and statistical significance of coefficients, goodness of fit and residual diagnostics, control variables, interaction terms and polynomial terms, variable selection and backward elimination, interpretation of logistic regression coefficients in terms of odds, multinomial regression
  • Correlation and causality: Difference between association and causal relationship, randomized controlled trials and observational studies, experimental design, confounding variables and their control via regression, limitations of causal conclusions drawn from observational data

Predictive Techniques and Data Representation


  • Fundamental concepts of predictive analysis: From describing observed data to predicting unobserved data. Training, validation, and test sets; cross-validation. Parameters and hyperparameters. Parametric and non-parametric methods. Linear and non-linear models
  • Generalization capability: Overfitting and underfitting as forms of model fitting to sample peculiarities rather than population regularities, bias and variance and their connection to model complexity, controlling complexity via variable selection and regularization
  • Linear regression as a predictive method: Evaluation metrics for regression problems (mean squared error and mean absolute error), evaluation on unobserved data, difference between a model that describes observed data well and one that predicts new data well
  • Classification techniques. Evaluating the performance of a classification model: Confusion matrix, precision, recall, and F1 score; behavior of metrics in the presence of imbalanced classes. K-Nearest Neighbors (KNN). Logistic and multinomial regression as predictive methods
  • The classifier as a decision rule on a score: Distinction between the quality of evidence provided by the data and the choice of the operational threshold. ROC curves and area under the curve (AUC) for studying the trade-off between the two types of errors as the threshold varies, interpretation of AUC in terms of observation ranking, threshold selection based on error costs, comparison between ROC curves and precision-recall curves in the presence of imbalanced classes
  • Generative classifiers: MAP decision rule as an application of Bayes' theorem, class-conditional density estimation, quadratic discriminant analysis (QDA) and linear discriminant analysis (LDA), Naive Bayes. Effect of different assumptions on the decision boundary and on the number of parameters to estimate
  • Features, representation functions, feature spaces, metrics
  • Clustering techniques: Definitions and K-Means
  • Fitting Gaussians to data, Maximum Likelihood, Gaussian Mixture Models
  • Non-parametric density estimation using Kernel Density Estimation
  • Principal Component Analysis (PCA)

-->

Textbook Information

Chapters from these books:

  1. Peck, Roxy, Chris Olsen, and Jay L. Devore. Introduction to statistics and data analysis. Cengage Learning, 2015.
  2. James, Gareth Gareth Michael. An introduction to statistical learning: with applications in Python, 2023.https://www.statlearning.com
  3. Bishop, Christopher M. "Machine Learning. Machine learning, 2006. https://www.microsoft.com/en-us/research/publication/pattern-recognition-machine-learning/
  4. Hernán, Miguel A., and James M. Robins. Causal inference, 2010. https://www.hsph.harvard.edu/miguel-hernan/causal-inference-book/
  5. Knaflic, Cole Nussbaumer. Storytelling with data: A data visualization guide for business professionals. John Wiley & Sons, 2025.

Teaching material shared by the teacher through Microsoft Teams (Team code: i87g4nb) and through the https://antoninofurnari.github.io/fadlecturenotes/ website.

Course Planning

 SubjectsText References
1Introduction to the course[1,2]; lecture notes
2Main Data Analysis Concepts [1]; lecture notes
3Descriptive Statistics and Graphical Representation of data[1]; lecture notes
4Uncertainty and data as the observation of random events[1,3]; lecture notes
5Probability for Data Analysis[1,3]; lecture notes
6Associations of Two Variables [1]; lecture notes
7Use of statistical inference in data analysis[1]; lecture notes
8Fundamentals of Causal Data Analysis[5]; lecture notes
9Storytelling with data[4]; lecture notes
10Predictive Data Analysis[3]; lecture notes
11Probabilistic Models for Classification[3]; lecture notes
12Clustering & Density estimation[3]; lecture notes
13Dimensionality Reduction and Principal Component Analysis[3]; lecture notes

Learning Assessment

Learning Assessment Procedures

The exam is divided into the following tests:

  1. A written exam designed to assess the student's theoretical understanding of the topics covered in the course, from both a theoretical and methodological perspective. The exam is graded on a scale of thirty.
  2. A project assigned by the instructor and carried out independently by the student, aimed at evaluating practical skills in data analysis and communication of results. The project is presented to the instructor through a presentation and graded on a scale of thirty.

Students with disabilities and/or DSA must contact the teacher, the CInAP representative of the DMI (Prof. Daniele) and CInAP well in advance of the exam date to communicate that they intend to take the exam using the appropriate compensatory measures.

Two written in itinere exams are scheduled during the course. Passing both tests grants exemption from the final written exam.

The final grade is obtained by means of a weighted average between the marks obtained in the two tests with weights of 40% for the written test and 60% for the project.

The assessment of learning can also be conducted remotely if the conditions require it.

The grading of each test is expressed on a scale of thirty points according to the following scheme:

Score 29-30 with honors

The student has a deep understanding of the concepts and techniques of data analysis. They can promptly analyze data analysis problems, identifying the most suitable data analysis techniques for the given problem independently and critically, and indicating the most suitable methodological practices for their application. They have excellent communication skills and language proficiency.

Score 26-28

The student has a good understanding of the concepts and techniques of data analysis. They can analyze data analysis problems, identifying appropriate data analysis techniques for the given problem and indicating suitable methodological practices for their application. They have good communication skills and language proficiency.

Score 22-25

The student has a fair knowledge of the concepts and techniques of data analysis, although it may be limited to the main topics. They can analyze data analysis problems, albeit not always in a linear manner, identifying suitable data analysis techniques for the given problem. They have fair communication skills and language proficiency.

Score 18-21

The student has minimal knowledge of the concepts and techniques of data analysis. They have limited ability to analyze data analysis problems. They have sufficient communication skills, although not always appropriate language proficiency.

Examination not passed

The student does not possess the minimum required knowledge of the main content of the course. Their ability to use specific language is very poor or nonexistent, and they are unable to independently apply the acquired knowledge.

Examples of frequently asked questions and / or exercises

The data analysis project is generally based on medium-large datasets obtainable on the internet.

Examples of typical exam questions:

  • Define the classification problem, discuss the differences with respect to the regression problem and give practical examples.
  • Explain the K-NN algorithm for classification. Discuss the effect of parameter K on algorithm performance. Give graphical examples of how the algorithm works and the effect of K.
  • Discuss evaluation measures for classification problems: accuracy, confusion matrix, precision, recall and F1 score. The pros and cons of the measures considered are discussed, also in relation to the characteristics of the test dataset.
  • Illustrate the main techniques useful for studying the correlation between variables.