What remains in what STT throws away
Speech recognition turns an utterance into text and discards the rest. The discarded side — rate, pause, rhythm, stress, discontinuity — also carries information. CASP is the research of turning that signal into quantitative indicators.
What disappears
The same sentence becomes a different event depending on how it is spoken. After passing through STT, that difference does not remain. A text-only system treats the two utterances as identical.
-
Pause
Where the break falls. A pause between sentences and a pause between words carry different meanings. What is observed is how an utterance changes before and after a long silence.
-
Rhythm
Whether speech flows evenly or breaks up. A collapse in rhythm is visible only in comparison with the speaker's own norm.
-
Energy
Loudness and stress distribution. Drops and surges are judged not by absolute value but within that person's usual range.
-
Transition
Change that occurs between one segment and the next. The indicator is not a value at a single point but the size and direction of movement.
Korean — how something is said changes the relationship
Honorific systems, sentence-ending processing, intonation, the placement of pauses. Korean is a language in which meaning and relational context shift sharply with the manner of utterance. The same vocabulary can become an entirely different statement.
The limit of English-centric models
Applying general-purpose overseas models directly can misflag a Korean speaker's normal speech patterns as anomalous signals. The range of normal variation differs by language.
The absence of reference data
There is not yet an industry-standard dataset defining the range of normal paralinguistic variation for Korean speakers. Without reference values, measuring an individual's change does not let you interpret it.
To measure an individual's change you must know their own norm; to judge whether that norm is unusual you must know the group's norm.
Two references
A personal baseline and a group distribution (Normal Range) play different roles. Merging the two causes an individual's distinctive traits to be misread as anomalous signals.
Personal baseline
Repeated measurement establishes that person's norm. Change is always a comparison against oneself.
Normal Range
The range of normal variation at the level of region, age, and sex. It lets an individual's indicators be read as a relative position within the group.
The Normal Range dataset for Korean speakers is currently being constructed, with the aim of establishing an interpretive standard that reflects regional, age, and sex distributions.
Current stage — engine complete, sampling under study
The anchor-point engine is complete. What remains under active research is how to assemble a sample suitable for it — that is, defining what to place at each end of the reference axis and constructing the Korean-speaker Normal Range that makes change measurable and interpretable. This is why CASP does not simply ingest arbitrary audio: the reference frame must be built before deviation can mean anything.
SQ · Systemizing Quotient
The systemizing-tendency scale proposed by Simon Baron-Cohen, referenced as the cognitive axis of Protocol v1.
Session-based repeated measurement
Recording is repeated by session, not a single measurement. A personal baseline is created only through repetition.
Anchor Point
A separate protocol that fixes the endpoints of the reference axis. What to place at each end must be set before an amount of change can be computed.
It structures where, by how much, and in what manner change occurred.
It speaks of an amount of variation — not cause, diagnosis, or intent. The aim is to redefine communication as a data structure that can be researched, validated, and applied, and it is oriented toward a Human-in-the-Loop approach.