Mentions & Social · Quantitative#225

Speaker-Idiolect Frequency Model

Every public figure has a stable verbal fingerprint — how often they use a given word per thousand words, and in which settings. Given a corpus of their past speeches, the probability that a word appears in a speech of known length and topic is a straightforward statistical estimate. Markets tend to price these on vibes about the news cycle instead. You build the corpus once and price every mention market from it.

What you need to run it

  • Transcript corpus per speaker, segmented by venue type and topic
  • Frequency model producing per-word appearance probability given speech length
  • Adjustment for current events that genuinely shift topic likelihood

Where this applies

Markets on Polymarket where speaker-idiolect frequency model is the natural play:

  • Will Trump say 'witch hunt' at his next rally?
  • Will the CEO mention 'AI' more than 10 times on the earnings call?
  • Will the candidate say 'healthcare' during the debate?

Capabilities this demands

Data ingestionModel / quantDomain knowledge

At a glance

CategoryQuantitative
MarketMentions & Social
Requirements3
CapabilitiesData ingestion, Model / quant, Domain knowledge
VenuePolymarket (CLOB, Polygon)

Build it

Related mentions & social strategies

This is documentation, not advice. Poly Research & Robotics publishes how these strategies work because the method should be checkable — not as a recommendation to trade them. See the full strategy database (297 strategies) or the data resources directory.
Join Discord