Ruixiang Tang

Teaching

Fall 2024 AI Topic Course: Trustworthy AI Foundations

Course Overview

This graduate topic course aims to give students a broader view of Trustworthy AI and focuses on understanding advanced techniques. The course covers key topics such as adversarial attacks and defenses, bias detection and mitigation, AI privacy, uncertainty estimation, and interpretable AI. Students will engage with the latest research, participate in discussions, and develop skills in critically analyzing and presenting complex material.

Prerequisites: This course will assume fundamental knowledge in AI and machine learning (e.g., 01:198:440 - Introduction to Artificial Intelligence, 16:198:536 - Machine Learning, or equivalent) and mathematical maturity (comfortable with linear algebra, probability, or equivalent). Students are expected to read and discuss research papers. Please contact the instructor if you have questions regarding whether your background is suitable for the course.

Grading

20% Quizzes

80% Final Project

We will have 5 short quizzes and one final project. The final project can be done individually or in groups of no more than 3. Presentations will take place during the last three weeks of the semester, followed by a Q&A session.

For the final project, students are required to choose a research paper on trustworthy AI, independently replicate its results, and critically analyze it to identify any limitations or potential areas for improvement. Students will then propose and develop a solution to address the identified issue or introduce a novel idea to enhance the research.

Course Schedule (tentative)

Week#TopicNotesRecommended Papers for Further Reading

Week 1

Introduction to Trustworthy AI

Overview of course objectives

Importance of Trustworthy AI

Key concepts and definitions

Trustworthy AI: A Computational Perspective

Week 2

Foundational Concepts in Deep Learning

Basic Knowledge of Deep Learning

Feedforward, Backpropagation

MLP, CNN, RNN,Transformer

Dive into Deep Learning

Week 3

Interpretable AI - Part 1

Introduction to Interpretable AI

Challenges of interpretability

Techniques for XAI

Techniques for Interpretable Machine Learning

"Why Should I Trust You?": Explaining the Predictions of Any Classifier

Grad-CAM: Visual Explanations from Deep Networks via Gradient-based Localization

Week 4

Interpretable AI - Part 2

Advanced methods for XAI

Case studies and applications

Evaluation of interpretability

Sanity Checks for Saliency Maps

A Benchmark for Interpretability Methods in Deep Neural Networks

Interpretation of Neural Networks is Fragile

Week 5

Bias Detection and Mitigation for AI Models- Part 1

Understanding bias in AI models

Sources of bias

Methods for detecting bias

Shortcut learning in deep neural networks

Gender Shades: Intersectional Accuracy Disparities in Commercial Gender Classification

ImageNet-trained CNNs are biased towards texture; increasing shape bias improves accuracy and robustness

Week 6

Bias Detection and Mitigation for AI Models - Part 2

Techniques for mitigating bias

Fairness in AI

Ethical considerations

Mitigating Gender Bias in Captioning Systems

Balanced Datasets Are Not Enough: Estimating and Mitigating Gender Bias in Deep Image Representations

Men Also Like Shopping: Reducing Gender Bias Amplification using Corpus-level Constraints

Week 7

Adversarial Attacks and Defenses for AI Models - Part 1

Introduction to adversarial attacks

Types of adversarial attacks

Case studies and examples

Explaining and Harnessing Adversarial Examples

Towards Deep Learning Models Resistant to Adversarial Attacks

Adversarial Examples Are Not Bugs, They Are Features

Trojaning attack on neural networks

Week 8

Adversarial Attacks and Defenses for AI Models - Part 2

Defense against adversarial attacks

Evaluation of defense strategies

Practical applications and challenges

Certified Adversarial Robustness via Randomized Smoothing

Automatic perturbation analysis for scalable certified robustness and beyond

Setting the Trap: Capturing and Defeating Backdoors in Pretrained Language Models through Honeypots

Week 9

AI Privacy

Introduction to AI privacy

Privacy Attack

Privacy-preserving techniques

Extracting Training Data from Large Language Models

Deep Learning with Differential Privacy

DP-OPT: Make Large Language Model Your Privacy-Preserving Prompt Engineer

Week 10

Trustworthy LLM - Part 1

Safety Alignment

Jailbreaking Attack and Defense

Multimodal Attack and Defense

Safety Alignment Should Be Made More Than Just a Few Tokens Deep

Assessing the Brittleness of Safety Alignment via Pruning and Low-Rank Modifications

Robust Prompt Optimization for Defending Language Models Against Jailbreaking Attacks

Week 11

Trustworthy LLM - Part 2

Hallucination Detection and Mitigation

Uncertainty Estimation

Large Language Models Can be Lazy Learners: Analyze Shortcuts in In-Context Learning

Extrinsic Hallucinations in LLMs

Detecting hallucinations in large language models using semantic entropy

To Believe or Not to Believe Your LLM

Week 12

Guest Lecture / Industry Speaker

Invited talk from a leading expert in Trustworthy AI

Discussion and Q&A session

Week 13

Student Presentations - Part 1

Student presentations

Discussion and feedback

Week 14

Student Presentations - Part 2

Student presentations

Discussion and feedback

Week 15

Student Presentations - Part 3

Student presentations

Discussion and feedback