Elective Course - 3nd Semester (Autumn 2nd year)
According to the curriculum, students must choose two elective courses during the third semester
Greek, English
Course Structure
Duration: 13 weeks
Format: 2-hour lectures + 1-hour Lab
Language: English
Assessment: 20% Weekly exercises, 40% Capstone/Project, 40% Exam, 10% Bonus Exercise
Course requirements
The primary objective of this course is to equip students with the theoretical knowledge and practical skills required to apply modern machine learning techniques to financial decision-making and data-driven problem solving. Students will learn how to develop, evaluate, and interpret predictive models using Python while understanding the mathematical principles that underpin contemporary machine learning algorithms.
This course provides a comprehensive introduction to machine learning with a strong emphasis on financial applications. It combines statistical learning concepts, data preprocessing, predictive modelling, model validation, and explainable artificial intelligence within a practical programming environment. Throughout the semester, students will work with real-world financial datasets to develop reproducible and industry-relevant machine learning solutions.
The course explores a broad range of financial applications, including:
Throughout the course, students will gain experience in building complete machine learning workflows, from data acquisition and preprocessing to model evaluation, interpretation, and communication of results.
Key Features
Best practices for reproducible research and machine learning workflows
Each weekly session consists of a 3-hour lecture designed to integrate theoretical concepts with practical financial applications.
The lecture is divided into two complementary components:
This integrated teaching approach enables students to connect theoretical knowledge with practical implementation while developing analytical and programming skills applicable to real-world financial problems.
Each lecture in part II is accompanied by a 1-hour practical laboratory session where students apply the concepts introduced during the lecture through guided programming exercises.
Laboratory activities include:
Every week, students will complete a practical exercise designed to reinforce the lecture material and progressively build the skills required for the capstone project.
Students are encouraged to responsibly use modern Artificial Intelligence tools (e.g., ChatGPT, Gemini, Claude, GitHub Copilot) to assist with programming, debugging, documentation, and understanding machine learning concepts. These tools are intended to enhance learning rather than replace it. Students remain fully responsible for understanding, explaining, and justifying all submitted code, modelling decisions, and analytical results.
Upon successful completion of this course, students will be able to:
Week 1: Python Foundations I
Aim: Develop the fundamental programming skills required for data analysis and machine learning using Python and Google Colab.
Topics: Introduction to Google Colab and Jupyter Notebooks; Python syntax and programming fundamentals; variables, data types, control structures, functions, file handling, virtual environments, and package management. Loading financial datasets, computing descriptive statistics, and producing introductory data visualisations.
Application: Introduction to financial datasets and exploratory analysis using Python.
Weekly Exercise 1: Develop a Python notebook implementing basic programming tasks, grouped descriptive statistics, and visualisations for a financial dataset.
Recommended Readings:
Laboratory Session 1: Google Colab setup, Python programming fundamentals, importing datasets, descriptive statistics, and introductory visualisation.
Week 2: Python Foundations II – Data Manipulation with NumPy and Pandas
Aim: Develop proficiency in manipulating and exploring structured financial data.
Topics: NumPy arrays; Pandas DataFrames; indexing, filtering, joins, group operations, data transformation, missing values, correlations, and exploratory data analysis.
Application: Exploratory analysis of a real-world credit risk dataset.
Weekly Exercise 2: Perform a complete exploratory data analysis and identify key data quality issues and statistical characteristics.
Recommended Readings:
Laboratory Session 2: Exploratory Data Analysis (EDA) using Pandas and visualization libraries.
Week 3: Data Cleaning and Preprocessing
Aim: Construct reliable and reproducible datasets suitable for predictive modelling.
Topics: Missing value treatment, duplicate detection, outlier identification, feature scaling, categorical encoding, preprocessing pipelines, and ColumnTransformer.
Application: Data preparation for credit risk modelling.
Weekly Exercise 3: Build a complete preprocessing pipeline and document the most significant data quality improvements.
Recommended Readings:
Laboratory Session 3: Building preprocessing pipelines using Scikit-learn.
Week 4: Feature Engineering for Financial Data
Aim: Develop meaningful predictive variables while avoiding information leakage.
Topics: Feature construction, financial ratios, interaction terms, nonlinear transformations, feature selection, leakage prevention, and domain-driven feature engineering. Special topic is the binning and woe.
Application: Credit risk and fraud detection.
Weekly Exercise 4: Design and evaluate 6–10 new predictive variables.
Recommended Readings:
Laboratory Session 4: Feature engineering techniques for financial datasets.
Week 5: Logistic Regression for Credit Risk Modelling
Aim: Develop an interpretable baseline classification model.
Topics: Logistic regression, class imbalance, class weighting, L1/L2/Elastic-Net regularisation, coefficient interpretation, probability calibration, and model diagnostics.
Mathematical Foundations: Odds and log-odds, maximum likelihood estimation, gradient optimisation, and regularisation penalties.
Application: Probability of Default (PD) modelling.
Weekly Exercise 5: Develop and compare regularised logistic regression models using cross-validation.
Recommended Readings:
Laboratory Session 5: Logistic Regression implementation for credit scoring.
Week 6: Model Evaluation and Performance Metrics
Aim: Evaluate predictive models using appropriate statistical and business performance measures.
Topics: Confusion matrix, Precision, Recall, F1-score, ROC-AUC, PR-AUC, KS Statistic, Gini coefficient, Log-Loss, Brier Score, Lift Charts, Gain Charts, and probability calibration.
Mathematical Foundations: ROC and Precision–Recall analysis, Negative Log-Likelihood, and the relationship between KS and Gini.
Application: Performance evaluation of credit risk models.
Weekly Exercise 6: Develop a reusable model evaluation toolkit for future assignments.
Recommended Readings:
Laboratory Session 6: Comprehensive evaluation of classification models.
Week 7: Decision Trees and Random Forests
Aim: Develop nonlinear predictive models and understand ensemble learning.
Topics: Decision Trees (CART), Gini impurity, entropy, pruning strategies, Random Forests, feature importance, partial dependence, and hyperparameter tuning.
Application: Credit risk prediction using ensemble methods.
Weekly Exercise 7: Optimise Random Forest models and compare predictive performance against Logistic Regression.
Recommended Readings:
Laboratory Session 7: Decision Trees and Random Forest implementation.
Week 8: Model Validation and Generalisation
Aim: Develop reliable models capable of generalising to unseen data.
Topics: Overfitting, underfitting, training-validation-test methodology, Stratified K-Fold Cross-Validation, nested cross-validation, learning curves, validation curves, and early stopping.
Application: Robust validation strategies for financial machine learning.
Weekly Exercise 8: Evaluate learning behaviour and diagnose model generalisation.
Recommended Readings:
Laboratory Session 8: Cross-validation and model diagnostics.
Week 9: Gradient Boosting Methods
Aim: Develop high-performance machine learning models for structured financial data.
Topics: Gradient Boosting, XGBoost, LightGBM, CatBoost, regularisation, hyperparameter optimisation, early stopping, feature importance, and model comparison.
Mathematical Foundations: Functional gradient descent and additive modelling.
Application: Advanced credit scoring and fraud detection.
Weekly Exercise 9: Optimise Gradient Boosting models using hyperparameter search.
Recommended Readings:
Laboratory Session 9: Building high-performance Gradient Boosting models.
Week 10: Alternative Machine Learning Algorithms
Aim: Compare complementary machine learning techniques for financial classification.
Topics: Support Vector Machines, k-Nearest Neighbours, Naïve Bayes, kernel methods, distance metrics, probabilistic classifiers, stacking method and optional anomaly detection using Isolation Forest.
Application: Benchmarking multiple classification techniques.
Weekly Exercise 10: Evaluate an alternative classifier and justify its suitability.
Recommended Readings:
Laboratory Session 10: Comparative evaluation of machine learning algorithms.
Week 11: Model Calibration and Explainable Artificial Intelligence
Aim: Improve model reliability and interpretability.
Topics: Probability calibration, Platt Scaling, Isotonic Regression, reliability diagrams, SHAP values, global and local explanations, and model documentation.
Application: Explainable credit scoring models.
Weekly Exercise 11: Produce an interpretable and calibrated Probability of Default model accompanied by a SHAP analysis and concise model documentation.
Recommended Readings:
Laboratory Session 11: Model explainability using SHAP.
Week 12: Introduction to Neural Networks
Aim: Introduce deep learning methods for structured financial datasets.
Topics: Artificial Neural Networks, multilayer perceptrons, activation functions, backpropagation, dropout, regularisation, early stopping, and model calibration.
Application: Credit risk and fraud detection using deep learning.
Optional Exercise: Compare the predictive performance of Neural Networks with the best traditional machine learning model developed during the course.
Recommended Readings:
Laboratory Session 12: Developing neural network models using TensorFlow/Keras.
Week 13: Capstone Project
Students will integrate the knowledge acquired throughout the semester by developing a complete machine learning solution using a real-world financial dataset.
The project should include:
Deliverables:
The following textbooks provide the theoretical and practical foundations of the course and are strongly recommended throughout the semester.
1. James, G., Witten, D., Hastie, T., Tibshirani, R., & Taylor, J. (2023).An Introduction to Statistical Learning with Applications in Python (2nd ed.). Springer.
2. Hastie, T., Tibshirani, R., & Friedman, J. (2009).The Elements of Statistical Learning (2nd ed.). Springer.
3. Géron, A. (2023).Hands-On Machine Learning with Scikit-Learn, Keras & TensorFlow (3rd ed.).
4. Goodfellow, I., Bengio, Y., & Courville, A. (2016).Deep Learning.
Students interested in advanced financial machine learning are encouraged to consult the following references.
Required Software
Students will complete all programming activities using Python and modern open-source data science libraries.
Required software includes:
Additional Software
Students may also use
Datasets
The course uses publicly available financial datasets throughout the semester. Typical datasets include
Additional datasets may be introduced during laboratory sessions and the capstone project
Academic Integrity
Academic honesty is an essential component of professional and scientific practice.
Students are expected to:
Any breach of the University's Academic Integrity Policy may result in disciplinary action in accordance with institutional regulations.
Responsible Use of Artificial Intelligence
Artificial Intelligence tools are becoming an integral part of modern data science and software development. Students are encouraged to use AI systems responsibly as learning assistants.
Acceptable uses include:
Students remain fully responsible for every line of submitted code, all modelling decisions, and the interpretation of results. During assessments, students may be required to explain or reproduce their work independently.
Collaboration Policy
Collaborative learning is encouraged throughout the course, particularly during laboratory sessions and class discussions.
The following guidelines apply:
Assessment Submission Policy
Students are responsible for submitting all coursework before the published deadlines.
Late submissions will normally incur a penalty in accordance with the University's assessment regulations unless an approved extension has been granted.
Requests for deadline extensions should be supported by appropriate documentation and submitted as early as possible.
Attendance and Participation
Regular attendance and active participation are strongly recommended.
Students are expected to
Consistent engagement throughout the semester will significantly improve learning outcomes and preparation for the capstone project.
Programming Standards
Students are expected to write clean, well-structured, and reproducible Python code.
Submitted notebooks should include
This course has been designed to combine rigorous academic foundations with practical machine learning skills that reflect current industry practice in finance and data analytics.
Its distinguishing characteristics include:
1. Practical Python Development
2. Complete Machine Learning Workflow
3. Financial Industry Applications
4. Explainable and Responsible AI
5. Industry-Relevant Programming Practices
6. Capstone Project
7. Research-Informed Teaching