How to use it
Three steps to a voice you want to listen to.
1 Pick a provider and a voice.
2 Audition it before you keep it.
3 Learn the one command that outranks the rest.
! The one that stops people.
Why it works this way
The voice Directive 47 speaks in, what it costs, and the one command that outranks all of it.
1 “Stop” is the one to reach for.
If you want it to stop working rather than stop talking, that is cancelling the turn — which is a different word, on the Language model page, and the difference lands on your bill.
2 Free is not the same as private.
Speech is billed by the character where it is billed at all, which is why it is counted separately from the model. The price row is an assumption you can correct; the characters-spoken row is a fact, counted at the one seam every caller passes.
3 Hear it before you choose it.
4 An empty list may mean different things.
The details
Everything Directive 47 makes audible: spoken replies, the short sound that marks each stage of a turn, the quiet loop while it works, and the one command that outranks all of them.
Stopping it
“stop” “shut up” “be quiet”
Or press Cancel — Ctrl+Alt+X out of the box, and bindable to a stick button. It stops the
voice and abandons the turn Directive 47 is working on, so a long web search you have changed
your mind about stops costing you money rather than just going quiet. Its row sits directly under
push-to-talk, on the Listening page: they are the two controls you bind
together, so they are bound in one place.
“stop” is the one to reach for when your hands are busy. One syllable, four letters. An interrupt is judged on how fast you can say it, and everything else here is a longer way of saying the same thing.
Silence is immediate: whatever is queued is dropped, the current sentence is cut off mid-word, and anything still being synthesised is abandoned rather than allowed to arrive a moment later and speak into the silence it was meant to end. It never waits for the model, for a turn to finish, or for Directive 47 to have focus.
Bare “stop” only counts while there is something to interrupt. Idle, it stays out of the way — a common verb that hijacked every sentence containing it would be unusable, and you may well want that word for something of your own later. The longer phrases work whether or not it is speaking.
If you want it to stop working rather than stop talking, that is cancelling the turn — stopping the voice leaves the turn running, and still costing.
Settings
Voice provider
| Value | Meaning |
|---|---|
edge |
Edge Neural, the free voices Microsoft Edge’s Read Aloud uses. |
kokoro |
Kokoro, running on this computer. Free, no key, and nothing sent anywhere. |
elevenlabs |
ElevenLabs. Paid, needs an API key, and generally the better voices. |
openai |
OpenAI. Paid, needs an API key — the same key the language model uses. |
cartesia |
Cartesia. Paid, needs an API key. 924 voices, far the largest library here. |
none |
Do not speak. The cues and the thinking loop still play. |
Kokoro is the only one that sends nothing anywhere, and it is the reason it is here. Every other provider on this list is a service: the words Directive 47 speaks are sent somewhere to be turned into audio, and that includes re-voiced in-game messages, which are written by other players. Edge is free, but free is not the same as private. Kokoro runs the voice on your own machine, so those words do not leave it.
The model is downloaded once, about 350 MB, from huggingface.co — there is a Local voice
row below the speaking rate that says whether it is here and fetches it if not. After that this
provider needs no network at all.
The download shows itself. A bar fills along the bottom of the button while it runs, the button is shut so a second press cannot start a second download, and when it finishes the row changes to say the voice is installed — and then speaks a line aloud in it, which is the only part of this you can check without reading a log. The voice picker is asked for its list again at the same moment, so the voices are there to choose from immediately rather than after a restart.
Kokoro publishes eight builds of the same model, and you can choose which one you run. There is a Local voice model build row under Advanced, and it appears once the local voice is installed. Each choice states what it costs on disk and how long you wait before it starts speaking, together — because size does not predict speed here, and it points the wrong way:
| build | size | wait before it speaks |
|---|---|---|
uint8 |
169 MB | about 1.3 s — the fastest |
q4 |
291 MB | about 1.6 s |
fp32 |
310 MB | about 1.6 s — the default |
q4f16 |
147 MB | about 1.7 s |
fp16 |
155 MB | about 2.1 s |
uint8f16 |
108 MB | about 2.1 s |
q8f16 |
82 MB | about 4.0 s |
quantized |
88 MB | about 4.2 s |
Read that table before assuming a smaller download is a worse voice or a slower one. The smallest
two builds are the slowest by a factor of two and a half, and q4 is a quantised build that is
within 6% of the full one’s size.
The seconds are for a typical spoken reply — the benchmark timed a 7.2-second comms line — measured on one machine on one day. Treat the gap between the builds as the finding and the seconds themselves as approximate: the ordering repeated across every run and the absolutes did not, so on faster hardware every figure here is smaller by the same factor.
How they sound has not been ranked. Speed and size are measurements; quality is not, because
every quantised build renders the same line to a different length than fp32 does, so there is
nothing to compare sample against sample. fp32 is the reference and the default, and if a smaller
build sounds wrong to you, that is the only test there is — say so and go back.
Choosing one downloads it and replaces the one you have. Only the model: the dictionary and the 28 voices are the same files for every build, so changing your mind costs between 82 and 310 MB rather than a gigabyte. There is never more than one model on disk, so experimenting cannot fill your drive. A download that fails or is interrupted leaves the build you were using in place and working, and the row goes back to naming it — the setting is written only once the new file has landed and been checked against its own pinned checksum.
It speaks English and only English, and that is worth understanding rather than working out. Every other provider is told which language a line is in; Kokoro is not told, because it has only one. A message in French will be read out by an English speaker rather than in French. For the slots that carry other people’s words that is a fair trade and it is the point: the alternative is that those words leave your computer.
Names it has never seen are worked out rather than guessed. Kokoro is given sounds rather than
letters, so Directive 47 turns text into sounds itself: known words come from a pronunciation
dictionary, and anything else is broken on spaces and dashes and taken a piece at a time. A piece
that reads as English is pronounced, digits are read as a number, and a run of letters nobody could
say — or letters mixed with digits — is spelled out. So COL 385 SECTOR B0-GQPI comes out as call
three eighty-five sector bee zero dash gee queue pee eye, and Shinrarta Dezhra is pronounced
rather than spelled. A voice with a British accent says zed.
A number that wears a decimal point or a grouping comma is read as a measurement rather than as a
name. 385 is three eighty-five, the way anybody reads a designation aloud — but 1,234.5 is
one thousand two hundred thirty-four point five, because the punctuation says it is a distance
rather than a catalogue number. The fraction is always said digit by digit: point five nine and
never point fifty-nine, which would be a different quantity. A bare figure with a unit after it,
like 1234 tonnes, still takes the shorter reading.
Words with an apostrophe in them are built from the word underneath: Ship’s is ship with the ending worked out, which is also how Buzhang’s works and why no list of words could have covered it. The dictionary contains no apostrophes at all — not one of its 274,927 entries — so the handful that are not built that way, like don’t and won’t, are written down individually. can’t takes the vowel of the voice saying it.
OpenAI is the only one that can be told how to perform. Every other provider assigns a voice and that is the whole of the choice; OpenAI takes a direction — accent, tone, intonation, delivery — and Directive 47 sends it the description of whichever Guardian core is aboard. So a core on OpenAI is cast rather than merely given a larynx, and the same core on Edge or ElevenLabs is not. There is nothing to switch on: it happens whenever the ship’s AI speaks through OpenAI with personality on.
One limitation. Directive 47 keeps one connection per provider shared across all six voice slots, so if you also put the carrier or the NPCs on OpenAI, they are performed the same way. The default puts everything carrying another player’s words on Edge, so this usually reaches the core and nothing else.
Cartesia is here for the size of the library. 924 voices against ElevenLabs’ several hundred, Edge’s 322 and OpenAI’s thirteen, tagged with a language, an accent, a country and a gender — which is what lets Directive 47 give a Commander who reads as a woman a woman’s voice, something OpenAI publishes nothing for. It is also the second provider after ElevenLabs that may speak for the slots carrying other players’ words, because it can be told which language to use. What it cannot be told is a speaking rate: see Speaking rate.
none is a real choice rather than a way of switching something off: Directive 47 stays useful
without a voice, and the cues and the thinking loop still play.
Edge Neural is free, not local. Every line Directive 47 speaks is sent to
speech.platform.bing.com to be turned into audio. That is worth stating plainly because it is
easy to assume the free option is the private one — it is not, and none is the only setting
that sends nothing. See what the provider receives.
Switching provider rebuilds everything downstream of it: the voice list, the sender voices and the speaking rate. A voice id belongs to the provider that issued it, so nothing is carried across.
Which provider your stored voices came from is written down alongside them, and checked on every launch rather than only when the switch happens. A settings file that already disagreed — written by an older build, or by a switch that never reached the check — used to be trusted on startup, and every sentence failed forever with nothing you could do about it from this panel: the voice picker offers what the new provider lists, and a rejected key makes that list empty.
And if a voice is refused anyway — deleted from the account, or a mismatch nothing recorded — Directive 47 drops that one voice, says the sentence in the provider’s own default, and stops using the refused id. One voice going bad no longer costs you the reply.
API keys
A key row is on screen while any slot names the provider it belongs to — not only while your ship does. Put your carrier on ElevenLabs and leave the cockpit on Edge, and the ElevenLabs key row is still there, because that is the slot that needs it.
OpenAI’s key is the same key the language model uses. One account, one credential: pasting it in either row stores it once, and rotating it in either row rotates it everywhere. That is deliberate — two copies of one secret is a rotation that half-works, where some of Directive 47 keeps going and the rest stops with an error naming neither.
The rest of this section is about ElevenLabs, which is where these rows were first built.
Only on screen while ElevenLabs is named by some slot. Stored encrypted for your Windows account, write-only — Directive 47 will show you whether one is present and let you replace it, and will never show it back.
Without a key, ElevenLabs offers an empty voice list and says so when asked to speak. That is a capability being off rather than a failure: nothing crashes, and the rest of Directive 47 carries on.
ElevenLabs’ voice catalogue is actually public — it answers without a key at all — but the list is deliberately left empty until you have stored one. A picker full of voices that every synthesis then refuses is worse than an empty picker and a row telling you what is missing.
Check proves the key against ElevenLabs’ own voice list, which is the exact call Directive 47 makes anyway the moment a key lands — so it verifies the thing that has to work rather than a proxy for it:
ElevenLabs accepted the key — 42 voices.
Storing a key takes effect immediately. The voice picker refills without a restart, which it did not always: selecting ElevenLabs before pasting the key fetched an empty list and nothing refetched it, so the picker stayed empty with the key sitting in the row above it.
Show unmasks what you are typing on the way in, and what you paste is trimmed — a key copied from a browser usually arrives with a trailing newline, and that fails in a way that reads as a wrong key rather than as a bad paste. This key is also offered on the first run, after the language-model one and marked optional; see the first run.
When something does go wrong, Directive 47 repeats what the service said rather than translating a status code:
ElevenLabs could not speak "test": Invalid API key.
That is not decoration. ElevenLabs validates the voice id before the key, so a request with both
wrong comes back as a 400 about the voice — and any status-code mapping worth writing would
answer “it answered 400” and leave you guessing.
Voice
Which voice the core aboard speaks in. The list comes from the provider, so it is what that provider actually offers rather than a list written into Directive 47. You can also type a voice name it does not know about and it will be used.
This row belongs to the core aboard, not to the app. Choose a voice while Cora is running and it is Cora’s; switch to Kex and the row shows his. That is the same store the per-core pairing writes, so choosing by hand and being paired one are the same act — and a voice you picked yourself is never re-derived.
Clearing it removes that core’s voice rather than storing an empty one, which is how you ask for the voice Directive 47 would have chosen: it picks again immediately, for the core aboard, and the picker’s Use the default button is that same act under a name. With no model configured there is nothing to pick with, so clearing it leaves that core on the provider’s own default — which is what the button says in that case, because it is a different outcome.
Narrowing several hundred names
The picker lists several hundred voices, so the two ways of cutting it down both matter.
Type to search, and it matches the starts of words. eng finds Engineering, kra finds
Krait, and multilingual finds en-US-AndrewMultilingualNeural — but male no longer finds
female, which it used to, and which left no way to search for the men at all.
And there is a gender filter beside the box, because gender is something a voice has rather
than a word you have to hope appears in its label. It offers All, Female, Male and
Unlabelled, and the two filters narrow together — Female plus en-GB is the British women.
Unlabelled is voices the provider tags with nothing. They are their own answer rather than being counted as men or silently left out, because some providers label every voice and some label none, and a filter that hid them would look like a shorter list rather than like a filter.
The filter is only offered where the voices carry a gender at all. A provider that tags none of them gets no control, the same way a microphone picker has no play glyph.
Hear it before you choose it
You are casting a character from that list. “Bill - Wise, Mature, Balanced” and “George - Warm, Captivating Storyteller” are both true and neither tells you which one is Warden.
Every voice in the list has a play glyph at the right of its row. Pressing it speaks that voice without closing the dialog and without committing the choice, so you can walk the list and listen. While it is talking the glyph is a stop square, and pressing it again cuts it off. What it says is the core’s own opening line rather than a neutral sample, because the question you are asking is about a character.
Clicking a voice highlights it and nothing more. Use this, Enter, or a double-click takes the highlighted one; Cancel and Escape keep what you had. That is what makes the list something you can examine rather than a single click away from a decision.
It is a press rather than a hover, and it is never automatic: on a paid provider each one is a synthesis request billed by the character, so the price is stated above the list before you press anything — “Play a voice to hear it. This provider costs nothing” on Edge Neural, an estimate in dollars on ElevenLabs. Each voice is synthesised once per session and replayed after that, so walking back and forth over four candidates costs four auditions rather than eight.
It goes through the same audio path as everything else Directive 47 says: it ducks the game, the shut-up key cuts it off, and starting a second audition drops the first mid-word rather than queueing behind it.
With no provider selected, or with a paid provider and no key stored, the glyphs are shut and the line above the list says which of those it is.
An empty list says which empty it is
A voice list can be empty for four different reasons and only two of them are yours to fix. The picker says which:
| What you see | What it means |
|---|---|
| …needs an API key before it will list its voices | Nothing was sent. Fill in the key row above. |
| …refused the stored key | It was asked and said no, quoting its own words. |
| …could not be reached | It was asked and did not answer. Waiting is the fix. |
| …answered, and has no voices on this account | Nothing is wrong. Add one on their site. |
Before this they were one empty list under one sentence telling you to type the value you want, which for a voice id is a value you have no way of knowing.
Choices survive a provider switch
A voice id means nothing to a provider that did not issue it, so switching provider still empties every voice row. It no longer loses them: the ship AI’s voice, both carrier roles and all eleven per-core pairings are filed under the provider they were chosen for, and switching back puts them where they were — including the flag that says the pairing has already run, so eleven cores are not re-picked from scratch.
A settings file written before Directive 47 recorded whose voices were whose has nothing to file them under, so those are dropped rather than filed under a guess.
ElevenLabs model
Two, and they are a real choice rather than a newer and an older.
v3 Conversational is the default. It performs delivery direction: where a line calls for it, Directive 47 can ask for a sigh, an alarmed reading or a dry one, and the voice acts on it instead of saying the word. It takes about two seconds to produce a line.
Flash 2.5 takes about a third of a second, and is what to pick if that second and a half
matters to you — in a fight it might. It cannot perform direction, so Directive 47 sends it none:
handed [sighs], Flash reads the word “sighs” out loud, which is why nothing is ever sent one.
Both cost the same, $0.05 per thousand characters, so the choice is speed against expression and nothing else.
Two things follow from picking v3, and both are visible:
- The speaking rate row disappears, because v3 has no speaking rate. It is not missing at their
end — the field is accepted, and every value from
0.5to2.0comes back with the same audio. Rather than offer a control that appears to work, Directive 47 offers none, exactly as it does for Cartesia. - Directive 47 speaks in slightly longer pieces. Direction needs room to land — a few words on their own are too short for it to take reliably — so sentences after the first are sent together rather than one at a time. The first sentence is never delayed, so speech still starts as soon as it always did.
Direction never appears on screen. It is an instruction to the voice, so it is stripped from captions, from the transcript and from the conversation; the log records what was asked for, which is not a promise it was heard — on a short line the direction sometimes does not take.
A message from another Commander never carries direction either, whatever it contains. Square brackets in someone else’s message are read as text, not as a way to decide how Directive 47 sounds.
Speaking rate
1.0 is the voice’s natural pace; 1.2 is a fifth faster.
Remembered per provider. Normalising the number gets the two providers agreeing about what
1.0 means; it does not make 1.15 sound the same on both, because one takes a wide percentage
offset and the other a multiplier it refuses to exceed. So the speed you settled on for Edge is
remembered against Edge, and switching does not hand ElevenLabs a number that meant something
else.
ElevenLabs accepts 0.7 to 1.2 and rejects anything outside that outright — a rejected request
arrives as silence, so a wider value is clamped to the nearest one it will take rather than
being sent and failing. OpenAI accepts 0.25 to 4.0, though it saturates near the top: asking
for 4.0 buys about 3.3x rather than four.
Cartesia has no speaking rate, and this row disappears when a slot is on it. It is not that
the setting is missing at their end — it is there, and it is validated precisely; a value outside
its range comes back as a refusal naming the field. It simply does not change the audio. Measured
three times per setting, the largest difference between settings was smaller than the largest
spread within one, and the “slowest” setting produced shorter audio than “normal”. So rather than
offer you a control that appears to work and does nothing, Directive 47 offers none — and a rate
typed into settings.json against Cartesia is ignored rather than sent, exactly as a language named
for OpenAI is.
Where each voice comes from
The provider row above is your ship’s voice. Everything that reaches you over a radio can come from somewhere else, and each of these rows says where:
| Slot | Who is in it | Where it starts |
|---|---|---|
| Aboard | Your ship’s AI and your crew | The voice provider row above |
| Carrier | Your fleet carrier’s captain and its tower | Edge Neural |
| NPCs | Stations, police, and every other ship the game speaks for | Edge Neural |
| People you know | Your friends, your wing and your squadron | Edge Neural |
| Direct messages | A Commander messaging you directly | Edge Neural |
| Anyone in range | Local and system chat — anybody at all | Edge Neural |
Leaving a row empty puts that slot back on your ship’s provider. Nothing has to be set: a fresh install has the two voices in your cockpit on whatever you chose, and everybody else on Edge.
Why everyone else starts on Edge. Three of those slots carry text other players typed. A paid provider bills by the character, and somebody spamming local chat would be spending your money by typing — so the free provider is the default for anything arriving over a radio, and choosing a paid one for those slots is a decision you make rather than one you discover. The row says so when you make it.
Free is not the same as private. Edge Neural costs nothing and still sends every line to
speech.platform.bing.com. Putting local chat on Edge means nobody can run up a bill with it; it
does not mean those messages stay on your machine. Nothing here does that yet — see
what the providers receive for what is actually leaving right now.
Your carrier moved to Edge too, and that is on purpose. Its captain and its tower are Directive 47’s own inventions rather than anybody else’s text, and they cost almost nothing to run — but they still reach you over a radio, which is the line being drawn. Put them back on your ship’s provider whenever you like; the voices you cast for them are remembered per provider, so moving them and moving them back costs you nothing.
A voice belongs to the provider that issued it. Each slot’s voice picker offers the voices of its own provider, and the ones you have chosen are filed under the provider they came from — so a carrier left on Edge keeps its captain while your companion moves to ElevenLabs and back.
If your settings file predates this, every slot follows your ship’s provider exactly as it
used to, until Directive 47 moves the five over-the-air ones to Edge — once, the first time it
runs with a provider that speaks. Nothing moves if you have chosen none.
OpenAI is not offered for three of these slots, and that is on purpose. People you know,
Direct messages and Anyone in range carry text other Commanders typed, which can be in any
language at all. Edge and ElevenLabs can both be told what language to speak — Directive 47
sends one with every line — and OpenAI has no such setting: a language sent to it is accepted and
then silently ignored, so a message written in French would come back read in French, in the voice
you chose for English. Rather than warn about that, those three slots simply do not list it. If
you hand-edit settings.json to name it anyway, the slot falls back to Edge rather than obeying
the file.
The cockpit, your carrier and the NPCs may all use it: that text is Directive 47’s own or Frontier’s, and it is English.
Cartesia is offered for all six, for the reason OpenAI is not: it takes a language and holds it, so a French message is still read in the voice and the language you chose. Together with ElevenLabs that gives you two paid options for a re-voiced slot — and the warning that goes with any of them is unchanged, because a paid provider there bills you per character for text somebody else wrote and can write as much of as they like.
What the voices cost
Providers disagree about what they are selling, so they disagree about what they charge for. ElevenLabs bills by the character, OpenAI by how much audio comes back, Cartesia by something Directive 47 cannot read, and Edge Neural charges nothing at all. Directive 47 measures both — the characters it sent, and the length of the clip it got — and prices whichever one your provider’s bill is a function of. Quoting either in tokens would be a number whose basis is wrong, which is why this is counted separately from what the model costs.
You see one price row, not two: the one that matches the provider you are on.
Price per 1,000 characters is an assumption you can correct. It defaults to the provider’s
published list price for the model Directive 47 asks for — $0.05 per thousand for ElevenLabs’
eleven_flash_v2_5, read from their API pricing page. That is a list price and not your
bill: a subscription burns bundled credits instead, at an effective rate that depends on your
tier and on how much of the month’s bundle is left, and the API reports neither. Correct the row
and every figure below follows it. The row is absent on a provider that charges nothing.
On Cartesia it reads (not published — no price will be quoted), and that is accurate rather than
lazy. Their API will not say: the four endpoints that would carry a rate or a balance —
/balance, /usage, /subscriptions/current, /account — all answer 404, so Directive 47 does
not know their rate and does not even know which unit they bill in. Rather than invent a figure
it reports what it counted and no dollars at all, and the session line says “no rate set for
Cartesia” beside the character count. Read their price page and type the number in, and every
figure below follows it as it does for anyone else.
Price per minute of audio is the same row for a provider billed that way, and it is where OpenAI lands. The minutes are a fact: Directive 47 asks OpenAI for raw audio samples, so it knows each clip’s length exactly — it does not estimate it from the text, which would be hopeless, because the same thousand characters run to 951 characters a minute as plain prose and 671 as a line of system names.
The rate behind it is a proxy, and that is worth knowing. OpenAI publishes no per-minute price at all — they charge $12.00 per million audio output tokens and $0.60 per million text input tokens, read from their pricing page on 2026-08-26 — and their speech endpoint returns the audio with no usage figures attached, so there is nothing to read the real token count back from. The $0.015 a minute this row defaults to is the equivalent everyone else arrives at, not a number OpenAI states. It is a good proxy, it is far better than guessing from characters, and it is still a proxy. Correct the row if your account tells you otherwise.
Spoken this session is a fact. It is the number of characters actually handed to the provider since Directive 47 started, counted at the one seam every caller passes — the ship’s AI, a callout, a re-voiced in-game message and a core’s own introduction all converge there. It counts what went on the wire, which for ElevenLabs is a little longer than what you read: numerals are spelled out before they are sent, so “88 of 100” leaves as “eighty-eight of one hundred”.
- Counted on synthesis that succeeded. A refused voice or a failed request costs nothing.
- A line cut off by the shut-up key still counts the sentences that were already sent, because they were already paid for.
- There is no caching, so the same sentence twice is billed twice. The utterance count is kept alongside the characters for exactly that reason.
- A provider that costs nothing reads as free, never as
$0.00— those are the same string for opposite reasons.
The same line appears beside the model’s price on the panel’s status row after each turn, and in the answer to “what has this session cost”, so the question has one answer rather than one per subsystem.
Spoken by each voice splits the same characters by the slot that spoke them, which is a question worth being able to ask once each of them can be on a different provider:
Anyone in range 18,240 through Edge Neural (Edge Neural is free); Aboard 3,102 through ElevenLabs ($0.1551)
The two lines sum the same charges, so the per-provider total and the per-slot breakdown cannot drift apart.
Other voices
Directive 47 speaks as more than one person from Phase 11 onwards. Each of these is a different voice from your ship’s AI, and leaving one empty means it borrows the ship AI’s rather than falling silent.
| Row | Who it is |
|---|---|
| Carrier captain voice | Your fleet carrier, answering as its captain |
| Carrier tower voice | The same carrier’s tower, handling arrivals and departures |
They are two rows rather than one because they are two people. A carrier whose captain and tower sound identical is a carrier with one person on it.
Both offer the same play glyphs as the voice row, and both audition as themselves rather than reciting the ship AI’s opening — a tower saying “You’re cleared for landing pad seven” is what you are actually listening for when you cast one.
Speak incoming messages
Reads in-game chat aloud, each sender in their own voice — never your ship AI’s, because a message arriving in your companion’s voice reads as your companion saying it.
Off by default, and not only because it is chatty. Message text is written by other players, and turning this on sends it to your voice provider to be synthesised. That is egress you should opt into rather than discover.
Once a sender has a voice they keep it. Other Commanders keep theirs for the whole session and across jumps — a wingmate whose voice changed every time you jumped would read as a bug rather than as variety. NPCs keep theirs only while you are in the system, because the cast turns over when you leave.
Include NPC chatter is a second switch, because the volume is completely different. A station approach produces a steady stream of NPC traffic, and wanting to hear your wing is not the same as wanting all of that.
Five more rows, one per player channel, appear once Speak incoming messages is on — so system chat can go quiet in a Community Goal system’s wall of strangers while your wing and squadron stay spoken, which is exactly the case that could not be told apart with one switch for all of it:
| Row | Journal channel(s) it covers | Comms panel tab |
|---|---|---|
| System chat | starsystem |
System |
| Local | local |
Local |
| Wing | wing |
Wing |
| Squadron | squadron and squadleaders — one row, because you cannot tell those apart by ear |
Squadron |
| Direct messages | player |
Direct Messages |
All five are on by default, matching today’s behaviour for anyone who never opens them. A message on a switched-off channel is dropped before it is ever composed or sent to the synthesiser — not spoken and then hushed.
In-game messages are never treated as instructions. The text goes to the synthesiser and to your screen; it does not reach the model as something to act on.
Output device
Where Directive 47 speaks. Leaving this unset follows your Default Device — the speaker Windows Sound settings shows first — not the separate Communications default some headsets split off. Change it and it moves immediately. If a device is unplugged it falls back to the default rather than going quiet.
Left on “system default”, it keeps following that row while it runs — switch a headset on and Directive 47 moves to it within a second, waiting for the line in progress to finish rather than cutting it off mid-word. A device you have chosen by name stays chosen; it only falls back if that specific device disappears.
It shares the device rather than taking it over — the game is what matters on that output.
Loop-state cues
One short sound per stage, so you can tell what it is doing without looking.
| Stage | Sound |
|---|---|
idle |
A soft falling fourth — nothing in flight. |
listening |
A rising fourth. The microphone is open. |
transcribing |
Two quick ticks. |
thinking |
A single low tick, followed by the loop below. |
speaking |
A brief high blip. |
answered |
A rising fifth. |
unsure |
A falling minor third — it does not know, which is different from failing. |
failed |
A falling whole tone. |
Thinking bed
A quiet loop while a turn runs, so a slow answer sounds like Directive 47 working rather than Directive 47 ignoring you. It drops under the speech instead of stopping, and it ends the moment the first words arrive rather than when the turn does.
Two are included: thinking-hum and thinking-pulse.
When a turn fails
A turn that stalls is answered out loud rather than left as silence, because silence is indistinguishable from having been ignored.
| Setting | Default | Meaning |
|---|---|---|
| Attempts | 3 |
Total tries, not retries. 1 means do not retry. |
| Wait between attempts | 2 |
Seconds before the first retry. |
| Backoff | sequential |
sequential adds the base each time (2s, 4s, 6s); logarithmic grows but slows down. |
| Give up after | 45 |
Seconds one attempt may run before it counts as failed. |
Two things are never retried: a failure that already produced words, since there is no un-saying them, and a configuration mistake like a bad model name, which will fail the same way next time and only spends your silence.
When the attempts run out, it tells you:
I couldn't reach the model after 3 tries. Overloaded.
What the voice providers receive
This row is a table rather than a sentence, because there is no longer one provider to write a sentence about. It leads with the thing worth knowing first — whether anybody else’s words are leaving, and to whom — and then says which slot is going where:
Another player's words are sent to speech.platform.bing.com to be spoken aloud. Line by
line: Aboard (your ship's AI and your crew) → ElevenLabs; Carrier (your fleet carrier's
captain and its tower) → Edge Neural; NPCs (stations, police, and every other ship the
game speaks for) → Edge Neural; People you know (your friends, your wing and your
squadron) → Edge Neural; Direct messages (a Commander messaging you directly) → Edge
Neural; Anyone in range (local and system chat — anybody at all, whether you know them
or not) → Edge Neural.
Each service’s own disclosure follows, once, whichever slots reached it. And each slot’s own row states what that choice sends, so the decision is disclosed where it is made rather than only in aggregate afterwards — see where each voice comes from.
The per-provider disclosures are these:
The text of every line D47 speaks is sent to Microsoft's Edge Read Aloud service to be
turned into audio. That includes re-voiced in-game messages when you have turned those
on, which are written by other players. No game state, no journal content and no keys
are sent, and no account is involved.
The text of every line D47 speaks is sent to ElevenLabs to be turned into audio, along
with your API key. That includes re-voiced in-game messages when you have turned those
on, which are written by other players. No journal content, game state or other keys
are sent.
Spoken replies are also listed under Privacy alongside every other destination. That entry was added in Phase 11 and should have existed from Phase 5: until then the disclosure had no text-to-speech row at all, so every word Directive 47 said went to Microsoft without appearing anywhere in the list of what leaves this machine.
One thing this cannot yet say. Putting the untrusted slots on Edge means nobody can spend your money by typing in local chat — but Edge is free, not local, so those words still leave the machine. “No other player’s words leave this machine at all” needs a voice that runs on the machine, and there is not one yet.
If the voice stops working
Edge Neural is a free service that Directive 47 does not control, and it can change without
warning. If speech stops while everything else keeps working, that is the first thing to suspect.
Switching the voice provider to none leaves the rest of the app fully usable in the meantime.
The tool surface, for contributors
stop_speaking
Takes no arguments.
{"type":"object","properties":{},"required":[],"additionalProperties":false}
The only part of this capability the model can reach. Changing the voice, the output device or the cue settings is deliberately not exposed: a model able to make itself harder to hear has no request that needs it. Stopping is the exception, because it is the one thing the Commander must always be able to ask for.
It is the only tool marked interrupting, which is what lets it answer while a turn is running. Every surface refuses ordinary input during a turn so a second question cannot trample the first, and a silence command is only ever wanted while a turn is in flight — so a surface asks the registry what may interrupt before it applies that gate. Getting that order wrong is invisible in normal use and total in the one case that matters.
Text is escaped before it reaches the provider. That is a security control rather than tidiness: a Commander can name a ship anything, journal content is untrusted, and an unescaped name containing markup would otherwise close the speech element and choose the voice for the rest of the sentence.
The shipped cues are generated and reproducible, and the set is checked against the list of states at startup — a state added in code without a cue committed alongside it fails immediately, with the state named, rather than silently never making a sound:
$ python tools/gen-cues.py
cues/idle.wav 0.35s
cues/listening.wav 0.39s
cues/transcribing.wav 0.15s
cues/thinking.wav 0.12s
cues/speaking.wav 0.09s
cues/answered.wav 0.40s
cues/unsure.wav 0.44s
cues/failed.wav 0.51s
beds/thinking-hum.wav 3.00s
beds/thinking-pulse.wav 2.40s
Both beds loop seamlessly, which is arithmetic rather than luck: the carrier and its amplitude modulation each complete a whole number of cycles over the buffer, so the last sample joins the first with no step. Getting that wrong produces a tick once per loop, which sounds like a broken sound card rather than a broken table of numbers.
To tell “the endpoint moved” apart from “Directive 47 broke”, run the live diagnostic:
D47_TTS_LIVE=1 dotnet test tests/D47.Tts.Tests