Fall 2024 AI Topic Course: Trustworthy AI Foundations
- Lectures: Wednesday 12:10-1:30 and Friday 2-3:20pm in Richard Weeks Hall on Busch campus room 208
- Instructor: Ruixiang Tang
- Office Hours: Friday 3-4:00pm, Hill Center room 416
Course Overview
This graduate topic course aims to give students a broader view of Trustworthy AI and focuses on understanding advanced techniques. The course covers key topics such as adversarial attacks and defenses, bias detection and mitigation, AI privacy, uncertainty estimation, and interpretable AI. Students will engage with the latest research, participate in discussions, and develop skills in critically analyzing and presenting complex material.
Prerequisites: This course will assume fundamental knowledge in AI and machine learning (e.g., 01:198:440 - Introduction to Artificial Intelligence, 16:198:536 - Machine Learning, or equivalent) and mathematical maturity (comfortable with linear algebra, probability, or equivalent). Students are expected to read and discuss research papers. Please contact the instructor if you have questions regarding whether your background is suitable for the course.
Grading
20% Quizzes
80% Final Project
We will have 5 short quizzes and one final project. The final project can be done individually or in groups of no more than 3. Presentations will take place during the last three weeks of the semester, followed by a Q&A session.
For the final project, students are required to choose a research paper on trustworthy AI, independently replicate its results, and critically analyze it to identify any limitations or potential areas for improvement. Students will then propose and develop a solution to address the identified issue or introduce a novel idea to enhance the research.
Course Schedule (tentative)
| Week# | Topic | Notes | Recommended Papers for Further Reading |
|---|---|---|---|
Week 1 | Introduction to Trustworthy AI | Overview of course objectives Importance of Trustworthy AI Key concepts and definitions | |
Week 2 | Foundational Concepts in Deep Learning | Basic Knowledge of Deep Learning Feedforward, Backpropagation MLP, CNN, RNN,Transformer | |
Week 3 | Interpretable AI - Part 1 | Introduction to Interpretable AI Challenges of interpretability Techniques for XAI | Techniques for Interpretable Machine Learning "Why Should I Trust You?": Explaining the Predictions of Any Classifier Grad-CAM: Visual Explanations from Deep Networks via Gradient-based Localization |
Week 4 | Interpretable AI - Part 2 | Advanced methods for XAI Case studies and applications Evaluation of interpretability | Sanity Checks for Saliency Maps A Benchmark for Interpretability Methods in Deep Neural Networks |
Week 5 | Bias Detection and Mitigation for AI Models- Part 1 | Understanding bias in AI models Sources of bias Methods for detecting bias | Shortcut learning in deep neural networks Gender Shades: Intersectional Accuracy Disparities in Commercial Gender Classification |
Week 6 | Bias Detection and Mitigation for AI Models - Part 2 | Techniques for mitigating bias Fairness in AI Ethical considerations | Mitigating Gender Bias in Captioning Systems Men Also Like Shopping: Reducing Gender Bias Amplification using Corpus-level Constraints |
Week 7 | Adversarial Attacks and Defenses for AI Models - Part 1 | Introduction to adversarial attacks Types of adversarial attacks Case studies and examples | Explaining and Harnessing Adversarial Examples Towards Deep Learning Models Resistant to Adversarial Attacks |
Week 8 | Adversarial Attacks and Defenses for AI Models - Part 2 | Defense against adversarial attacks Evaluation of defense strategies Practical applications and challenges | Certified Adversarial Robustness via Randomized Smoothing Automatic perturbation analysis for scalable certified robustness and beyond Setting the Trap: Capturing and Defeating Backdoors in Pretrained Language Models through Honeypots |
Week 9 | AI Privacy | Introduction to AI privacy Privacy Attack Privacy-preserving techniques | Extracting Training Data from Large Language Models Deep Learning with Differential Privacy DP-OPT: Make Large Language Model Your Privacy-Preserving Prompt Engineer |
Week 10 | Trustworthy LLM - Part 1 | Safety Alignment Jailbreaking Attack and Defense Multimodal Attack and Defense | Safety Alignment Should Be Made More Than Just a Few Tokens Deep Assessing the Brittleness of Safety Alignment via Pruning and Low-Rank Modifications Robust Prompt Optimization for Defending Language Models Against Jailbreaking Attacks |
Week 11 | Trustworthy LLM - Part 2 | Hallucination Detection and Mitigation Uncertainty Estimation | Large Language Models Can be Lazy Learners: Analyze Shortcuts in In-Context Learning Extrinsic Hallucinations in LLMs Detecting hallucinations in large language models using semantic entropy |
Week 12 | Guest Lecture / Industry Speaker | Invited talk from a leading expert in Trustworthy AI Discussion and Q&A session | |
Week 13 | Student Presentations - Part 1 | Student presentations Discussion and feedback | |
Week 14 | Student Presentations - Part 2 | Student presentations Discussion and feedback | |
Week 15 | Student Presentations - Part 3 | Student presentations Discussion and feedback |