Communications·production

IAM.Speech

Transcription, diarization, QA scoring and conversation reports.

Key capabilities

  • ASR and speaker diarization
  • Quality assurance with keyword and intent analysis
  • Reports, fact tables and result export

Step-by-step guides

Use cases

Start with the outcome: open a guide, prepare prerequisites, follow the steps and verify the success signals.

01 Upload a recording and get a diarized transcriptUtterances are separated by speaker and synchronized with timestamps.
Audience
QA analyst
Outcome
Utterances are separated by speaker and synchronized with timestamps.

Before you start

  • Audio/video file or signed URL
  • Recording language
  • Expected speaker count

Steps

  1. Create an analysis jobUpload the file, select language and enable diarization.
  2. Configure speakersSet min/max speakers when known from the conversation type.
  3. Wait for ReadyProcessing must complete ingest, ASR, diarization and alignment without failed stages.
  4. Review uncertain segmentsFilter low-confidence and speaker-change segments and listen at the timestamp.
Verify the result

Timestamps open the correct audio and speaker labels remain consistent.

If it does not work

Poor diarization requires checking channel mix/overlap rather than replacing all labels manually.

02 Configure QA criteria and export a reportCalls are scored by one scorecard version and findings link to utterances.
Audience
Quality manager
Outcome
Calls are scored by one scorecard version and findings link to utterances.

Before you start

  • QA scorecard
  • Test conversation set
  • Corrective-action owners

Steps

  1. Create the scorecardAdd criteria, weights, critical failures and allowed N/A.
  2. Validate on a reference setCompare auto score with manual labels and tune thresholds.
  3. Run a batchPin scorecard version and period; do not mix versions in one report.
  4. Review and exportOpen evidence for critical failures, then export CSV/XLSX.
Verify the result

Every score has criterion id, scorecard version and evidence timestamp.

If it does not work

For sudden score changes, compare scorecard version and sample composition before changing the model.

Application sections

Open detailed manual

Every application screen has a separate page with controls, safe example values, CLI/API alternatives and status-specific recovery steps.

Open detailed manual →

Role in the ecosystem

IAM.Speech processes audio and recorded communications, turning them into transcriptions, speaker segments, facts and quality control reports. Source could be a file upload, API, or a write from IAM.Comm/IAM.Voice.

Processing pipeline

ingest → normalize → VAD/ASR → diarization → QA/extraction → report

Each stage publishes state and maintains a connection to the source. Partial the result is clearly marked: the absence of diarization should not look like confirmed assignment of cues to specific people.

Possibilities

  • timestamps, punctuation and domain vocabulary;
  • speaker diarization and manual role correction;
  • keywords, intents, required phrases and QA score;
  • structured facts and report export;
  • batch API and reprocessing with a new configuration.

Quality and reproducibility

Metrics are divided by language, channel, noise and number of speakers. Model version, the processing profile and source checksum are stored next to the result. Reprocessing does not silently overwrite the previous version.

Safety

Retention of the source audio and text is set separately. No access to recording follows automatically from access to the aggregated report. Export and deletion leave an audit receipt.