A 3d environment for pedagogically testing Dutch NT2 conversations.
For the preset environment on SpeechLab Windows laptop, view User Guide here: https://uva.works.surf.nl/s/x59qJDGzCSBjMTS
Classroom_Scene_LLM from Assets/Scenes/Assets/Data/PiperModelsUpdate prompt for Convo Agent under LLMAgent chat settings. Refer to prompts from here: https://github.com/arneeichholtz/nt2-conv/tree/main/prompts
In play mode, move the player around the environment with arrowkeys or 'WASD'. Move towards the conversational avatar to trigger 'Convo Mode', which will lock the camera on the character and make the UI appear.
There are two modes of input: writing responses as text, or recording voice to turn speech to text (with microphone icon). On enter, the response is submitted to the LLM, and a reply will show on screen.
To exit 'Convo Mode', press 'esc'. This will bring you back to 'Gameplay Mode'. To enter 'Convo Mode' again, player mmust walk into the conversational avatar.
in the case that unity crashes while running, make sure that the models are properly attached as seen in the screenshot for the different gameobjects and components, expecially for SpeechToText, TextToSpeech, and TextToSpeech > RunPiper game objects
Hierarchy of game objects

The interactions are set up between two modes: gameplay and convomode. This has been implemented through a state machine architecture, with CameraPositions storing world positions for GameplayView and ConvoView.

![]()
![]()

LLMAgent, handles chat settings for LLM.

SpeechToText

TextToSpeech

TextToSpeech > RunPiper (nl)

Main Camera, has a camera controller, which animates between GameplayView and ConvoView

Main Camera > ConvoUI > Canvas The settings for the UI components can be found here.
StatusDisplay shows the status message for the different controls while in ConvoMode.Entry > TextInput is where text input is received, more specifically Text Area.
Record object contains the icon image for recording, as well as UI handling for when icon is clicked, to start/stop recording for speech to text.DialogueDisplay contains the UI components for responses received from the LLMAgent.
Speak object contains the icon image and button event handling for text to speech feature.FrequencyBandRenderer is a feature for when recording audio, to give visual feedback on frequency that is being recorded.LLM implementation via LLMUnity (v3.0.3) package: https://github.com/undreamai/LLMUnity Some code for TTS/STT adapted from: https://github.com/danielbierwirth/Inference-Whisper-Piper-Unity
A 3d environment for pedagogically testing Dutch NT2 conversations.
For the preset environment on SpeechLab Windows laptop, view User Guide here: https://uva.works.surf.nl/s/x59qJDGzCSBjMTS
Classroom_Scene_LLM from Assets/Scenes/Assets/Data/PiperModelsUpdate prompt for Convo Agent under LLMAgent chat settings. Refer to prompts from here: https://github.com/arneeichholtz/nt2-conv/tree/main/prompts
In play mode, move the player around the environment with arrowkeys or 'WASD'. Move towards the conversational avatar to trigger 'Convo Mode', which will lock the camera on the character and make the UI appear.
There are two modes of input: writing responses as text, or recording voice to turn speech to text (with microphone icon). On enter, the response is submitted to the LLM, and a reply will show on screen.
To exit 'Convo Mode', press 'esc'. This will bring you back to 'Gameplay Mode'. To enter 'Convo Mode' again, player mmust walk into the conversational avatar.
in the case that unity crashes while running, make sure that the models are properly attached as seen in the screenshot for the different gameobjects and components, expecially for SpeechToText, TextToSpeech, and TextToSpeech > RunPiper game objects
Hierarchy of game objects

The interactions are set up between two modes: gameplay and convomode. This has been implemented through a state machine architecture, with CameraPositions storing world positions for GameplayView and ConvoView.

![]()
![]()

LLMAgent, handles chat settings for LLM.

SpeechToText

TextToSpeech

TextToSpeech > RunPiper (nl)

Main Camera, has a camera controller, which animates between GameplayView and ConvoView

Main Camera > ConvoUI > Canvas The settings for the UI components can be found here.
StatusDisplay shows the status message for the different controls while in ConvoMode.Entry > TextInput is where text input is received, more specifically Text Area.
Record object contains the icon image for recording, as well as UI handling for when icon is clicked, to start/stop recording for speech to text.DialogueDisplay contains the UI components for responses received from the LLMAgent.
Speak object contains the icon image and button event handling for text to speech feature.FrequencyBandRenderer is a feature for when recording audio, to give visual feedback on frequency that is being recorded.LLM implementation via LLMUnity (v3.0.3) package: https://github.com/undreamai/LLMUnity Some code for TTS/STT adapted from: https://github.com/danielbierwirth/Inference-Whisper-Piper-Unity