booper (aka babble) is a language model that starts from random weights and learns only from what consenting people send it on Discord. This page is the whole deal, in plain words. The short version: nothing of yours is stored until you explicitly say yes, you can take it all back at any time, and your data is never sold.
If you opt in with !babble accept, the messages you address to the bot —
@mentions, replies to it, and DMs — are stored as plain text. Replies that start with the
correction marker are also filed as corrections. If you additionally run !babble all
in a channel, everything you say in that channel is collected too; it never
widens collection for anyone else, and !babble pings turns it back off. That is the
complete list. The bot does not scan channels, read history, or collect anything from people who
haven't opted in.
Collected text is used to train the model and is published in a public HuggingFace dataset that
anyone can download. In the published data you appear only as a salted hash (like
u_9f2c…), never as your Discord id or username. The salt never leaves the machine the
bot runs on, so the hash cannot be reversed back to you.
Your Discord id and username never leave the machine the bot runs on. Messages not addressed to
the bot are never stored (unless you turned on !babble all for that channel). Server
names, channel names, avatars, presence, and member lists are not collected at all. If you never
opt in, the bot still replies to you, but nothing from the exchange is stored or trained on.
Data collected through Discord is never sold, rented, licensed for advertising, or traded — full stop — and no such sale would ever happen without the express consent of Discord itself, which we do not intend to seek. No data used to train booper is ever sold. By the same token, booper is not a paid service and will stay free for as long as its corpus contains data collected through Discord, in keeping with Discord's terms on the resale of platform data.
!babble forget deletes everything of yours — corpus rows, correction pairs, and
your opt-in — immediately, and takes effect on the very next message. Your text is then excluded
from every future training run and every future published version of the dataset. To be honest
about the one limit: copies of the public dataset that other people already downloaded are out of
our reach. !babble consent shows you what scopes you've granted at any time, and if
you want to see what's stored about you, ask (see contact below) and we'll show you.
Data lives on a single machine controlled by the operator, not in a third-party cloud, and only the operator has access. Consent records and the id-to-hash mapping are kept only as long as they are needed to enforce your choices. There is no analytics, no tracking, and no third-party processor beyond HuggingFace hosting the already-anonymised public dataset.
booper is only for people old enough to use Discord under Discord's own terms (13, or higher
where local law says so). We do not knowingly collect text from anyone below that age; if we learn
we have, it gets deleted the same way !babble forget deletes anything else.
booper operates under the Discord Developer Terms of Service and Developer Policy. Everything above — explicit opt-in, no sale of platform data, honoring deletion, collecting the minimum — is how we hold ourselves to them.
If what we collect or publish ever materially changes, the bot re-asks you before any new collection happens — a policy change never silently widens your old yes. Questions, complaints, or data requests: open an issue at github.com/frgmt0/babble.