Do speech language models actually tell voices apart, or do they just go by the words?
And if not, how do we teach them to trust the voice, not just what is being said?
Dongwook Lee
I am an M.S. student in the Interdisciplinary Program in Artificial Intelligence (IPAI) at Seoul National University, advised by Sungroh Yoon. Previously, I received my B.S. from Seoul National University. My research interests center on speech-language models for natural spoken interaction, with a focus on real-time, human-like agents that can coordinate, interrupt, and continue conversations naturally.
Email: dwsmart32@snu.ac.kr / GitHub / LinkedIn
Consecutive translation waits for the sentence to end, so it comes out smooth.
Simultaneous translation cannot wait, so by nature it comes out in choppy pieces.
But there is never only one way to say the same thing. Can we use that freedom to translate as you speak, without the choppiness?
What happens when someone nearby tries to hijack the conversation?
How can we build voice assistants that stay robust to third-party interruptions?
Can offline RL generalize to unseen states by reusing the pieces of states it has already seen?
What if an unfamiliar state is not entirely new, but a new composition of familiar anchors and changes?
Researcher in speech and language intelligence
Lab Intern in reinforcement learning
B.S. in Naval Architecture and Ocean Engineering
B.S. in Computer Science
Exchange Student in Engineering College