Term Project
One of the course objectives is to learn how to pursue a research problem and communicate your findings with others. You will develop those skills in this class by doing a term project. This term, the project runs in two tracks—pick one:
- Track 1 (Research): Pursue an open research problem of your own choice.
- Track 2 (Reproducing prior work): Implement state-of-the-art techniques on an assigned topic and contribute them to a shared evaluation platform.
Teams and Sign-up (both tracks)
- A team can have up to 3 students. Given the class size, the course cannot accommodate solo term projects.
- Please use this Google Sheet [link] to form a team and sign up for a track (you need an ONID to access this sheet).
- [Deadline] 09/30. If your team is not on the sheet by then, your final score will decrease by 2%.
- Every team presents three times (see Presentations), regardless of the track.
Track 1: Research
You are encouraged to choose any topic relevant to computer security, but your research problem should yet to be explored in the literature. You are welcome to select a topic connected to your current research: for example, if your research area is reinforcement learning, you can do a term project on some aspects of security and privacy in reinforcement learning.
Identifying a research problem is part of the work in this track. Please start the topic search as early as possible—teams that are still searching at Checkpoint 1 tend to run out of time.
Track 2: Reproducing Prior Work
If you have a hard time finding a topic of your own, choose this track. You do not search for a topic; instead, you implement recent work on an assigned topic and evaluate it under a common setup.
Topic
Preventing and tracing the distillation of large language models (LLMs). Anyone with query access to a deployed LLM can distill it into a smaller model that inherits much of its capability. This term, teams in this track implement the state-of-the-art techniques that address the two sides of this problem:
- Prevention: making a deployed model hard to distill (e.g., perturbing or restricting what the model exposes, detecting extraction-style query patterns).
- Tracing: making distillation detectable after the fact (e.g., watermarks and fingerprints that survive into a student model, and provenance tests on a suspect model).
[Note] The specific methods to implement, and the models and datasets to run them on, are assigned by the instructor—you do not choose them. The authoritative list lives in the repository README; the repository and the per-team assignments are shared on Canvas after sign-up closes.
How this track works
The goal of this track is a working evaluation platform: a single repository in which each method is implemented, run on the assigned models and datasets, and reported under a common protocol, so that the results are comparable across methods and reproducible by someone else.
- The instructor sets up and shares the repository with the teams that sign up for this track.
- Your team branches off the repository—you never commit to the main branch directly.
- You follow the README instructions to lay out your folder and code: directory layout, interfaces, configuration files, and how results and logs are written.
- You implement your assigned method(s) and run the evaluation on the assigned models and datasets.
- You write your report on those evaluation results, using the same report template as Track 1.
[Note] This track is managed strictly. The point of a shared platform is that every contribution looks the same and runs the same way, so the conventions are not suggestions:
- Your code MUST follow the structure and interfaces specified in the README. Code that does not fit the layout will be sent back for revision.
- Your results MUST be reproducible from what you commit: pinned configurations, fixed random seeds, and the logs your runs produced.
- Your branch is reviewed before it is merged. Expect revision requests, and leave time for them.
- Run only the assigned models and datasets. If you believe something else is needed, ask first.
In exchange, this track spares you the effort of searching for a topic, and you end the term with a concrete, reviewed implementation of a recent technique—which is a good deal more portable than a slide deck.
Presentations
Your team should present the progress
three times; I have the following expectations:
Checkpoint Presentation 1 (on 10/12)
- Do a literature review and explain the prior work relevant to your project.
- [Track 1] Identify a research problem (or your project scope), and clearly articulate how your problem is different from the prior work.
- [Track 2] Explain the assigned papers in detail, and state how you plan to map each method onto the repository structure.
- Provide your next steps.
Checkpoint Presentation 2 (on 11/09)
- Design a set of experiments and obtain preliminary results.
- [Track 1] You must implement a minimal (key) functionality to demonstrate the feasibility of your idea.
- [Track 2] You must show at least the key result of one assigned method, produced by code running on your branch.
- Provide your next steps.
Final Presentations (on 12/02)
- Deliver your final results (based on the goals you set in Checkpoint 1).
- [Track 2] Your branch should be ready for review: code in the required layout, and results reproducible from your committed configurations.
- In-class presentation and a report [Note: the report is due 12/09, on Canvas].
Grading Policy (Evaluations)
Your term project will be evaluated with the following scheme (50% in total); the scheme is the same for both tracks:
- 15%: Checkpoint Presentation 1
- 15%: Checkpoint Presentation 2
- 20%: Final Presentation and Write-up (Report template)
- -2%: No team on the Google Sheet by 09/30.