Course Details

Distributed systems and big data

MF0651

Course
Distributed systems and big data
Code
MF0651
Academic Year
2023/2024
Curriculum Year
2023/2024
Degree Programme
ARTIFICIAL INTELLIGENCE AND DIGITAL INNOVATION
Curriculum
000 - 000-GENERICO
Course coordinator
Credits
12
Lecture Hours
96
Scientific Disciplinary Sector (SSD)
INF/01 - Computer Science
Course Type
Single-subject learning activity
Course Delivery
OPZ - Opzionale
Year
1
Teaching period
Secondo Semestre
Campus
ALESSANDRIA
Teaching language
Italian
Course Contents
Principles and current technical and methodological solutions for the design and management of systems, architectures, and distributed applications for Big Data analysis.
Reference Texts
For the first part, there is no specific reference textbook. The teaching material consists of scientific articles, technical documentation and other material provided by the lecturer.
However, the following textbooks cover most of the topics covered in the first part of the course:
* M. van Steen and A.S. Tanenbaum, "Distributed Systems, 4th ed.," 2023 (freely downloadable from the website: https://www.distributed-systems.net).
* M. Kleppman, "Designing Data-Intensive Applications: The Big Ideas Behind Reliable, Scalable, and Maintainable Systems," O'Reilly, 2017.
* J. S. Damji, B. Wenig, T. Das, and D. Lee, "Learning Spark: Lightning-Fast Data Analytics, 2/E," O'Reilly, 2020.

For the second part, the textbook can be downloaded from here: https://iris.uniupo.it/handle/11579/140999.

For the third part, there is no specific reference textbook. The teaching material consists of scientific articles, technical documentation and other material provided by the lecturer.
Learning Outcomes
The course aims to illustrate the principles and current technical and methodological solutions for the design and management of systems, architectures, and distributed applications for Big Data analysis. In particular, the course addresses the various problems concerning the acquisition and processing of Big Data, by presenting the main solutions (both hardware and software) that have been proposed for their management. The course also introduces the Rust programming language and the peculiarities that make it suitable for Big Data processing. The course includes hands-on exercises in order to be able to combine the methodological and technological aspects seen in the classroom.

* Knowledge:
At the end of the course, the student will have learned the principles and the main technical and methodological solutions for the design and management of systems, architectures, and distributed applications for Big Data analysis as well as advanced programming concepts especially relating to static analysis (through types) of memory.

* Competences and skills:
At the end of the course, the student will have learned the knowledge to achieve the educational objectives, in the way measured by the examination grade. In particular, the student will be able to autonomously investigate topics related to systems and architectures for Big Data, and to use this knowledge to critically evaluate existing systems and applications as well as to tackle new problems, and to evaluate the appropriateness of their implementability in Rust.
Prerequisites
Fundamentals of programming, database and operating systems.
Teaching Methods
Class lectures and hands-on lab lectures.
Additional Information
There are no mid-term exams.
Assessment Methods
Written exam and practical assignments.
In the written exam, the student must demonstrate that (s)he has learned the concepts and methodologies presented during the course.
In the practical assignments, the student must demonstrate that (s)he is familiar with the frameworks, platforms and tools presented during the course.

Detailed Syllabus
The course consists of three parts.
First part (6 CFU):
- Introduction to Big Data: motivations, principles, issues and challenges
- Data acquisition systems.
- Distributed file systems and data stores.
- Cluster resource management systems.
- Batch processing systems
- Stream processing systems.
Second part (3 CFU):
- Introduction to Cloud Computing.
- Theory and practice on Google Cloud Platform.
- Theory and practice on Amazon Web Services.
- Theory and practice on OpenStack.
Third part (3 CFU):
Introduction to Rust programming:
- Automatic memory management: ownership, borrowing and lifetime of references.
- Primitive and structured data types.
- Error handling and functional programming patterns.
- Generic types and traits
- Multithreading

Expected Learning Outcomes
* Knowledge and understanding:
At the end of the course, the student will have acquired the methodological knowledge to independently investigate topics related to systems and architectures for Big Data, and to use this knowledge to evaluate existing systems as well as to tackle new problems.

*Applying knowledge and understanding:
At the end of the course, the student will have learned the methodologies for the design and development of systems and applications for Big Data. Specifically, the student will be able to design and develop distributed, scalable, safe and efficient batch processing and data stream processing applications, using the main open-source frameworks for the ingestion, processing and storage of Big Data.

* Making judgments:
At the end of the course, the student will be able to independently identify the most suitable solutions to create distributed systems and applications for Big Data analysis and to evaluate both the architectural and implementation choices, and the performance of existing systems and applications for Big Data.

* Communication Skills:
At the end of the course, the student will have acquired the most suitable fluency in terminology related to systems and applications for Big Data analysis, will be able to present the architecture of a system for Big Data with appropriate technical terms and language and to argue critically about the various alternatives both at a system and application level.

* Learning Skills:
At the end of the course, the student will have acquired the methodological knowledge to independently investigate topics related to systems and architectures for Big Data, and to use this knowledge to evaluate existing systems, to tackle new problems as well as to evaluate advantages and disadvantages of choosing programming languages for the development.
Last update:09-09-2026 00:14:31