SayList

Building a Spoken-Task Corpus

A named set of spoken phrases and expected tasks turned extraction quality from a successful demo into a measurable regression problem.

The early prototype quickly encountered the difficult part of voice-to-task software: recognizing words is not the same as identifying obligations. One passage can contain several actions, connective phrases, deadlines, places, and fragments that should not become tasks.

By October 2025, development had moved from isolated examples to a named set of 25 phrases with expected results. The 7 October commit added audio fixtures, a gold-suite test, export checks, and a substantial revision of the rule-based extractor.

That choice gave later work a stable baseline. Dates, trigger phrases, duplication, and sentence splitting could be changed against a corpus instead of judged only from one successful demonstration.

The corpus did not solve natural language, and its audio and live-testing material remains outside the public export because publication rights and personal-data status require owner review.

← Back to SayList