Course Details

DATA AND PREDICTIVE ANALYTICS

EC0137

Course
DATA AND PREDICTIVE ANALYTICS
Code
EC0137
Academic Year
2026/2027
Curriculum Year
2025/2026
Degree Programme
MANAGEMENT, ECONOMICS AND FINANCE
Curriculum
A18 - Marketing and Operations Management
Course coordinator
Lecturers
Credits
6
Lecture Hours
45
Scientific Disciplinary Sector (SSD)
SECS-S/01 - Statistics
Course Type
Single-subject learning activity
Course Delivery
OBB - Obbligatoria
Year
2
Teaching period
Primo Semestre
Campus
NOVARA
Teaching language
Italian
Course Contents
The course presents the statistical methodology for the managment and the quantitative marketing with the support of ad hoc software.
Reference Texts
Some selected chapters in:

G. James, D. Witten, T. Hastie, R. Tibshirani (2021). An Introduction to Statistical Learning with Applications in R, 2nd edition, Springer.

Useful references:
P. H. Franses , R. Paap (2010) Quantitative Models in Marketing Research. Cambridge University Press.
Learning Outcomes
The course introduces the main statistical tools for management and marketing with the aid of an ad hoc software. The aim is to increase the knowledge on the main multivariate exploratory and predictive statistical techniques, to develop analytical skills through the use of these techniques and the corresponding IT tools and develop a critical capacity in the use of the same.
Prerequisites
Contents of basic Statistics and of basic Mathematics in economics (see E0252, E0362 and EA007).
Teaching Methods
The course offers two alternative learning tracks, structured to meet students’ varying attendance needs and educational goals.

Track A – Integrated Instruction (In-Person)
This track is ideal for students who can commit to regular attendance and wish to acquire both theoretical and practical skills. It consists of:
- Lectures: Theoretical presentation of statistical methodologies and their properties.
- Practical exercises: Sessions using specialized software for data analysis.
- Case study analysis: Critical discussion of the methods applied and interpretation of results guided by the instructor.
- Self-study: Supplementary and assessment activities supported by digital materials available on the DIR platform.
The lecture-based and interactive components are integrated into the course of study, accounting for approximately 2/3 and 1/3 of the total workload, respectively.

Track B – Independent Study (Theoretical)
This track is designed for non-attending students or for those who wish to focus exclusively on the conceptual and theoretical aspects of the subject, excluding the technical components of programming and software use. It consists of:
- Individual study: Independent learning based on reference texts and the recommended bibliography.
- Dedicated guidelines: Provision of specific instructions and detailed outlines on which chapters of the textbook to study in depth.
The self-study component is central.
Additional Information
Students with physical disabilities, Learning Disabilities or Special Education Needs can request specific services and tools via the Staff Sviluppo e Coordinamento Carriere e Servizi alle Studentesse e agli Studenti, consulting the University webpage: https://www.uniupo.it/en/services/servicesstudents-physical-or-learning-disabilities Students with disabilities, learning disabilities or special education needs, once they have contacted the University Staff, can refer to the tutor in charge of the course to define the examination modalities, concerning academic aspects.
Assessment Methods
The evaluation is based on an oral exam.

Track A
The exam consists of
- theoretical questions to assess knowledge of concepts and mastery of the terminology;
- short numerical exercises to test the skills acquired in the use of computational algorithms and specialized software;
- problem-solving or structured case studies based on data interpretation, critical discussion of the techniques used, and analysis of the results, aimed at evaluating the candidate’s ability to conduct statistical analysis independently.

Track B
The exam consists of
- theoretical questions to assess knowledge of basic concepts and mastery of the terminology;
- short numerical exercises to evaluate understanding of the basic mechanisms of statistical techniques;
- an explanation of the main algorithmic mechanisms underlying the primary methodologies.

Detailed Syllabus
The program varies depending on the track chosen by the student.

Track A
- Introduction to Statistical Learning: prediction and inference problems. Supervised and Unsupervised Statistical Learning.
- Introduction to the R software and the use of R Commander.
- The simple linear regression model. Meaning and OLS estimation of parameters. Confidence intervals and significance tests.
- The multivariate linear regression model: parameter estimation and significance tests, variance decomposition, and the F-test. Variable selection: cross-validation (CV) and predictive statistics (AIC, BIC, Cp). Leverage and high-impact points. Use of dummy variables for categorical predictors. Models with interactions. Multicollinearity: definition, effects, and identification (VIF). Normality of residuals: interpreting the QQ-plot. Multiplicative models and their linearization.
- Using R Commander features to comprehensively perform a linear regression analysis.
- Classical conjoint analysis. Experiment design: definition of factors, levels, and stimuli. Full and fractional factorial designs. Utility assessment: Likert scales. Model, partial stimuli, and assessment of the relative importance of factors.
- Supervised classification. Bayesian decision rule and selection of discriminant variables. Logistic model. Odds ratios and their relationship to posterior probabilities. Classification diagnostics: confusion matrix and ROC curve.
- Use of generalized linear models in R for supervised classification.
- Unsupervised Statistical Learning: clustering methods. Partitioning methods (k-means) and hierarchical clustering (linkage methods). Dendrogram and its interpretation. Using R Commander functions to perform clustering.
- Dimension reduction methods. Principal component analysis: definition and interpretation. Using R to obtain scores and the biplot.
- Nonparametric methods: CART. Splitting rules (RSS, entropy, Gini index). Pruning a tree. The R function CART().
- Introduction to neural networks: basic concepts of neural networks (layered architecture, simple examples of fully connected networks). The problem of model training: backpropagation.

Track B
- Introduction to Statistical Learning: prediction and inference problems. Supervised and Unsupervised Statistical Learning.
- The simple linear regression model. Meaning and OLS estimation of parameters. Confidence intervals and significance tests.
- The multivariate linear regression model: parameter estimation and significance tests, variance decomposition, and the F-test. Variable selection: cross-validation (CV) and predictive statistics (AIC, BIC, Cp). Leverage and high-impact points. Use of dummy variables for categorical predictors. Models with interactions. Multicollinearity: definition, effects, and detection (VIF). Normality of residuals: interpreting the QQ-plot. Multiplicative models and their linearization.
- Supervised classification. Bayesian decision rule and selection of discriminant variables. Logistic model. Odds ratios and their relationship to posterior probabilities. Classification diagnostics: confusion matrix and ROC curve.
- Unsupervised Statistical Learning: clustering methods. Partitioning methods (k-means) and hierarchical clustering (linkage methods). Dendrogram and its interpretation.
- Dimension reduction methods. Principal component analysis: definition and interpretation. Biplot.
- Nonparametric methods: CART. Splitting rules (RSS, entropy, Gini index). Tree pruning.
- Local regression: general principle and differences from linear regression.
- Generalized additive models. General aspects and estimation techniques.

Expected Learning Outcomes
The expected outcomes vary depending on the track chosen.

Track A
Fragmentary knowledge or knowledge containing substantial errors, an inability to correctly set up and solve exercises or problems, a lack of ability to use the software, and inadequate scientific language will result in a failing grade.
Essential and predominantly descriptive knowledge, with concepts applied through the solution of basic exercises; the ability to solve simple problems—even with some uncertainties—but with a comprehensible explanation, along with the use of the software’s basic functions, ensures a passing grade.
Complete mastery of the content, combined with a high degree of autonomy in formulating and solving problems; effective analysis and synthesis with relevant connections and a rigorous presentation allow students to achieve excellent grades.

Track B
A limited understanding of the main methodological tools, substantial errors in defining approaches, and inadequate scientific language lead to a failing grade.
Predominantly descriptive knowledge—albeit with some uncertainties but presented in an understandable manner—and the ability to apply concepts to basic exercises
ensure a passing grade.
Complete mastery of the content, combined with effective analytical and synthesis skills, relevant connections, and a rigorous presentation, enables students to achieve excellent grades.
Last update:09-09-2026 00:14:31