Repository navigation
feat: Social coordination domain pack - #43
Merged
Merged
Conversation
added 4 commits
October 2, 2026 01:35
A social platform, and the coordinated behaviour inside it: accounts,
topics and devices; a follow graph; an organic stream of posts, reshares
and replies; and labelled campaigns, some inauthentic and some not.
The social domain has been a GraphSchema preset since 0.5 and nothing
more — entities, a topology, one latent factor as its only ground truth,
and no tests at all. The fraud pack has a temporal process, a labelled
pattern catalogue, a measured hardness dial and a scoring harness. This
gives the social side the same treatment, as a domain of its own rather
than by changing `social`, because the graph is a different one.
Organic decoys are the point
----------------------------
Laundering has a weak legitimate twin. Coordinated behaviour has a
strong one: a fandom reacting to a release, a city reacting to an
earthquake, and a paid amplification ring are structurally the same
thing — a burst of near-simultaneous, near-identical activity from a
tightly connected group.
So a dataset whose only labelled structures are inauthentic is worse
than useless. It rewards any detector that fires on synchrony, and that
is exactly the detector that ruins real platforms by suspending fan
clubs. `fandom_burst`, `breaking_news` and `mutual_follow_community` are
generated by the same machinery as the campaigns, labelled
is_coordinated=False, and counted as false positives when flagged. Each
is matched to its twin on the properties that are *not* the point,
chiefly posting volume: a five-post fandom against a one-post copypasta
ring is separable by count alone, which is not the distinction a
detector should have to get right.
Eight inauthentic playbooks, including `follow_farm` specifically
because an event-only detector cannot see it — its members hardly post,
so `recall_by_playbook` exposes a one-signal detector.
Measured, not asserted
----------------------
`tradecraft` is an input to a measurement. Mean per-playbook
best-single-feature AUC falls 0.984 / 0.852 / 0.789 across low, medium
and high, and the evidence that still works broadens from identity and
timing to structure and content.
`examples/coordination_detectors.py` scores five rules a platform would
try first. The conclusion matches the fraud pack's Cypher cookbook:
single-signal detectors do badly. Best F1 anywhere is 0.499, at the
easiest setting. The co-occurrence rule goes from 0.843 precision with
no organic false positives at `low` to 0.118 precision with 15% of
organic accounts flagged at `medium`. `fresh + active` scores zero at
`low` — the setting where every campaign account *is* fresh — because
low tradecraft strips their cover activity, so they are not "active": a
conjunction failing where each half is strongest.
No label leakage, enforced
--------------------------
No node attribute names the answer, and a test asserts that no single
feature separates any playbook perfectly, because a perfectly separating
feature means a label leaked into the graph. It caught two during
development:
- account handover removed *all* of an account's organic activity,
leaving reshare_count at exactly zero and AUC 0.976 at every
tradecraft level. Real revived accounts have a quiet history and a
normal present, so dormancy is now a timestamp and activity after it
is kept.
- every sockpuppet cluster was separable at AUC 1.000 by device
sharing, because organic sharing was always exactly two accounts.
Shared tablets, cafés, NAT and fingerprint collisions put unrelated
accounts behind one identifier, so a few devices now serve a crowd.
Two more fixed while measuring: polars `unique(keep="first")` does not
maintain order, so which duplicate follow survived varied between runs
and broke reproducibility for a fixed seed; and an account in two
overlapping campaigns had its creation moved after the earlier
campaign's events, so creation is clamped to an account's first event.
Realism of the organic platform, at scale 0.0006 seed 7: 27% silent
accounts (real platforms are mostly lurkers, and without them "low
activity" is uninformative), event Gini 0.65, follower Gini 0.46,
reciprocity 50%, clustering 0.42, diurnal peak-to-trough 18.7.
Known limitation: follower Gini is lower than a real platform's, because
the tail here is bounded by follows_per_account while real inequality
comes from accounts with millions of followers. Posts carry a
template_id and no text; the playbooks are coordination shapes from
public platform-manipulation research, not operational recipes.
43 new tests, suite at 310. Docs in docs/domains/coordination.md.
The pyg sink already existed and is mostly domain-agnostic, but it
hardcoded the fraud pack's truth schema in three places, so a
coordination dataset failed on `ColumnNotFoundError: is_fraud`.
The label rules are now a convention rather than one pack's column
names: a truth frame keyed by `<entity>_id` with a single boolean column
labels those entities. `accounts.is_fraud` and `accounts.is_coordinated`
both work, and a third pack following the same shape gets labels without
touching this module. Requiring *exactly* one boolean means the
convention fails loudly instead of guessing between two.
Three bugs found while generalising:
- Matching a label frame on values alone picked `campaigns.topic`,
which holds real Topic ids next to a boolean, and labelled every
topic. The id column now has to end in `_id` as well as match a node
table's ids.
- Latent factors were detected by excluding fraud's frame names
(`patterns`, `accounts`, `transactions`). A frame with an `*_id`
column is labelling entities whatever it is called, which is the
property to test.
- Foreign keys were excluded from node features but not from edge
attributes, so the coordination pack's interaction edges one-hot
encoded the topic they are about into 48 columns — thousands at
scale. Identifiers are excluded too: a standardised `tx_id` or
`template_id` is a meaningless axis and the information it stands for
is in the structure.
What a coordination dataset now exports: Account with `y`
(coordinated), `decoy` (organic) and stratified splits; `POSTED`,
`RESHARED` and `REPLIED` with `y` and `edge_time`; `community` as a
latent tensor rather than a feature; `FOLLOWS` and `USES` unlabelled,
because the truth says nothing about them.
examples/coordination_pyg_baseline.py trains a heterogeneous GraphSAGE
against a logistic regression on account features alone, at all three
tradecraft levels. Coordination is a better fit for a GNN than fraud is,
because a campaign is defined by who acts with whom, so the features-only
control is genuinely weak.
tradecraft model AUC AP organic flagged
low features only 0.932 0.536 -
low + graph 0.964 0.785 -
medium features only 0.680 0.101 4.0%
medium + graph 0.791 0.243 12.0%
high features only 0.590 0.050 2.6%
high + graph 0.696 0.082 7.7%
The graph helps at every level and average precision more than doubles
at medium. It also flags three times as many organic accounts, because
it buys its power by learning "tightly connected group acting together"
and a fan club is exactly that. Measuring AUC alone gives "graphs win",
which is true and incomplete, and that gap is what the organic decoys
exist to make visible. The organic share is measured at a fixed
operating point — each model asked for as many accounts as there are
real campaign members — so models with different score distributions
compare fairly.
9 new tests, suite at 320. The 10 existing pyg tests now run rather than
skip, with torch installed locally.
The pyg job's end-to-end step only covered the fraud pack, which is the one whose truth schema the sink used to be written around. The coordination pack uses different column names (is_coordinated, event_id), so this is what actually checks that the label rules are a convention and not fraud's schema.
denironyx
force-pushed
the
social-coordination
branch
from
October 3, 2026 14:10
a0d8f47 to
7ccacf1
Compare
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
No description provided.