Skip to content

Project 3: Topic drift in patient–doctor chat

Mentors: To be announced

Problem: Patient–doctor conversations shift from symptom-gathering to diagnosis at some point, but that pivot isn't marked in the transcript. There's no way to locate it without reading the whole exchange.

Context: Built on MTS-Dialog (CC-BY-4.0) patient–doctor dialogues — each transcript's dialogue column is a real, expert-authored Doctor:/Patient: multi-turn exchange, the actual turn-to-turn structure this project traces, unlike the single question/answer pairs in Symptom2Disease (used instead in Project 7) — anchored to ML4LLM Ch.3 · proj9: Sequential word cosine similarity (helper).

Goals: Where in a conversation does embedding similarity between consecutive turns drop, i.e., where does the doctor pivot from symptom-gathering to diagnosis?

Deliverables: A notebook that embeds each conversational turn, computes sequential cosine similarity turn-to-turn, and plots the similarity trace across a conversation. Sharp drops mark topic pivots, compared across many conversations to see if pivots cluster around a predictable turn number.

Showcase: TBD

References:

← Back to all Basic Science projects