Differences

This shows you the differences between two versions of the page.

Link to this comparison view

Both sides previous revision Previous revision
Next revision
Previous revision
dcc:itsol:whisper [2025/12/17 08:20] – added beta release warning giuliodcc:itsol:whisper [2026/09/25 10:27] (current) – alba
Line 1: Line 1:
 {{indexmenu_n>5}} {{indexmenu_n>5}}
 +
 ====== Whisper Guide ====== ====== Whisper Guide ======
  
 ===== Attention! ===== ===== Attention! =====
  
-**The use of Whisper is at the user's own responsibility. The setup illustrated in this guide makes use of an instance of Whisper that is run locally and that does not send data outside of the local environment. Please make sure to handle your data correctly. If you have any doubts about how to do so, please start by following the steps on this page of our wiki: [[dcc:itsol:whisper:datamanage]].**+<WRAP round important>
  
-**If you have any questions on Whisper that this guide does not answer, please feel free to send us a message at [[dcc@rug.nl|dcc@rug.nl]].**+The use of Whisper is at the user's own responsibility. The setup illustrated in this guide uses a locally run instance of Whisper that does not send data outside the local environment. **Please make sure to handle your data correctly.** If you have any doubts about how to do so, start by following the steps in [[dcc:itsol:whisper:datamanage|Data management safety measures]].
  
-**News Item (16-12-2025):** We have released a beta version of the Whisper interface with the option to add diarization (speaker recognition) to the transcription/translation job. This version of the interface is not final yet, so please keep in mind that not everything might work the way you want it to. We will complete this new interface in January, but in the meantime, feel free to test it out using the default parameters.+If you have any questions on Whisper that this guide does not answer, send us a message at <dcc@rug.nl>. 
 + 
 +</WRAP>
  
 ===== Introduction ===== ===== Introduction =====
  
-This guide takes you through the steps to set up a series of folders and a script to run speech-to-text transcription on the University of Groningen infrastructure (for UG staff and students) based on the [[https://openai.com/research/whisper|OpenAI Whisper automatic speech recognition (ASR) model]] running on the [[https://iris.service.rug.nl/tas/public/ssp/content/detail/service?unid=0d51dd1aa44f4cdcb4949f1702d1829f|Hábrók High Performance Computing]] (HPC) cluster.+This guide takes you through the steps to set up a series of folders and a script to run speech-to-text transcription on the University of Groningen infrastructure (for UG staff and students) based on the [[https://openai.com/research/whisper|OpenAI Whisper automatic speech recognition (ASR) model]] running on the [[https://iris.service.rug.nl/tas/public/ssp/content/detail/service?unid=0d51dd1aa44f4cdcb4949f1702d1829f|Hábrók High Performance Computing]] (HPC) cluster. 
 + 
 +The process of transcribing spoken audio to text is usually a very time-consuming manual process. The UG offers a licensed version of [[https://www.audiotranskription.de/en/f4transkript/|F4 Transkript]] on the University Workplace as an aid for manual transcription, but doesn't offer automatic speech recognition software. 
 + 
 +This guide is offered by the DCC to help researchers process their research data as efficiently as possible, while optimizing data protection (keeping their audio files on UG storage instead of sending them to cloud services). The app is the work of three CIT teams: the DCC came up with the concept, the Data Science team developed an early command-line version, and the HPC team builds and maintains the app you run from the portal. If you wish to read more on the detailed functionalities of Whisper, refer to the [[https://github.com/openai/whisper|manual in their Git repository]]. 
 + 
 +Should you have any further questions on the use or initial setup of Whisper on Hábrók HPC, contact the DCC at <dcc@rug.nl>. 
 + 
 +===== What Whisper can do ===== 
 + 
 +Whisper turns speech in audio and video recordings into text. On Hábrók, it runs as a batch job that you start from the web portal, so you do not need to install anything or use the command line. 
 + 
 +The app can: 
 + 
 +  * transcribe speech in its original language; 
 +  * translate speech into English; 
 +  * align the transcript to the audio for more precise timestamps; and 
 +  * label who is speaking, through optional speaker diarization. 
 + 
 +The output is produced automatically and can contain errors. Check names, numbers, specialist terms, and quotations against the recording before you rely on the result. 
 + 
 +===== How it works ===== 
 + 
 +The guide follows four steps, one page each. Each page links on to the next. 
 + 
 +  - [[dcc:itsol:whisper:setup|Before you start]] — request an account, create the input and output directories, and upload your recordings. 
 +  - [[dcc:itsol:whisper:running|Run a Whisper job]] — open the app, fill in the form, and submit and monitor the job. 
 +  - [[dcc:itsol:whisper:results|Find and review the results]] — the output files, the job log, and what to check before you rely on a transcript. 
 +  - [[dcc:itsol:whisper:datamanage|Data management safety measures]] — what to remove from Hábrók once you have your transcripts. 
 + 
 +[[dcc:itsol:whisper:troubleshooting|Troubleshooting]] is there for when a job fails or the output is wrong; it also says where to get help and answers common questions. 
 + 
 +The clip below shows the whole process, from upload to finished transcript.
  
-The process of transcribing spoken audio to text is usually a very time-consuming manual process. The UG offers a licensed version of [[https://www.audiotranskription.de/en/f4transkript/|F4 Transkript]] on the University Workplace as an aid for manual transcription, but doesn't offer automatic speech recognition software.+{{ :dcc:itsol:whisper:end_to_end.webm?800x450 |}}
  
-This guide is offered by the DCC to help researchers process their research data as efficiently as possible, while optimizing data protection (keeping their audio files on UG storage instead of sending them to cloud services). For technical aspects, the service is supported by the Data Science and HPC team of the CIT. If you wish to read more on the detailed functionalities of Whisper, please refer to the [[https://github.com/openai/whisper|manual in their Git repository]].+For your first job, use one short recording. It finishes quickly and lets you check the settings and the output before you commit a whole collection to a run.
  
-Should you have any further questions on the use or initial setup of Whisper on Hábrók HPC, please contact the DCC at [[dcc@rug.nl|dcc@rug.nl]].+<WRAP rightalign>[[dcc:itsol:whisper:setup|Next: Before You Start →]]</WRAP>
  
-[[dcc:itsol:whisper:setup| → Move to the next step]]