# How a view is counted, and why the rules are what they are

This is the most important document in the repository. Everything else here is
a story website; this is the part that decides who gets paid.

Read it before changing anything in `server/lib/views.js`, and read it before
relaxing a threshold because an author complained.

---

## The problem, stated plainly

This platform pays novelists per view and earns from advertising. Those two facts
together create a specific, predictable failure:

**Paying per view gives every author a direct financial incentive to buy
traffic.** Bot traffic costs a few dollars per hundred thousand hits. If a view
is worth $0.0015 and a hundred thousand bot hits cost $5, an author who can get
those hits counted turns $5 into $150. That is not a hypothetical; it is the
economics of every per-view payout scheme ever built.

**The damage is not the payout.** Overpaying one author by $150 is annoying.
The real exposure is that the same fake traffic loads the ad slots. Google's
invalid-traffic detection is looking at the ad requests, not at our payout
ledger, and its remedy is not "we will not pay for those impressions" — it is
account termination. One author's bot farm can end the ad revenue for every
author on the site, permanently, with no appeal that anyone has found reliable.

So the design target is not "count views accurately". It is:

> Never pay for, and never serve ads against, traffic we could not defend to
> Google if asked.

That is deliberately a stricter bar, and it means we knowingly discard some
real reads. The asymmetry is the whole point: a lost genuine view costs an
author a fraction of a cent. A tolerated fake one can cost everybody
everything.

---

## The shape of the mechanism

A page load counts for nothing. Nothing at all.

```
  reader opens a chapter
        │
        ▼
  GET /api/v1/stories/:slug/chapters/:n
        │  server issues a SIGNED READ TOKEN
        │  (chapter id, required dwell, issue time, address hash)
        ▼
  reader reads. the page collects scroll depth, interaction
  count, and the gaps between interaction events
        │
        │  after the required dwell has actually elapsed
        ▼
  POST /api/v1/read/heartbeat  { readToken, dwell, scroll, interactions, gaps }
        │
        ▼
  qualify() — a pure function, no database, no clock of its own
        │
        ├── fails ──► read_events row, qualified = false, with every reason
        │
        └── passes ─► read_events row, qualified = true
                      + qualified_views row  ← this is the row that is money
```

Both tables are written **every time**. `read_events` is the audit trail;
`qualified_views` is the payable subset. The ratio between them is the single
most useful fraud signal the site has, and it only exists because the rejected
events are kept.

---

## The rules

Every rule below is enforced in `qualify()` in `server/lib/views.js` and pinned
by a test in `server/test/views.test.mjs`. The tests are written as the attacks
they defend against, because that is how each rule will be re-argued when
somebody wants it relaxed.

### 1. A view requires an explicit heartbeat after a minimum dwell time

The dwell floor is derived from the chapter's word count:

```
required = clamp(words / 600 wpm × 60s, 15s, 180s)
```

- **600 wpm** is roughly double a fast adult reader. The number is not a model
  of reading; it is a bound a human plausibly clears and a script has to
  actually wait out. Lower and genuine skimmers are excluded. Higher and a bot
  farm just adds a `sleep`, which costs them almost nothing.
- **15 second floor.** Without it, a 40-word "interlude" chapter would count
  instantly — which is precisely the thing a fraudulent author would publish
  two hundred of. The floor is what makes chapter-count padding pointless.
- **180 second ceiling.** A 12,000-word chapter must not demand twenty minutes
  of unbroken attention, or honest long-form readers stop being counted.

**The dwell that matters is the age of the server-issued token, not the number
the client sends.** The client's claimed `dwellMs` is used for exactly one
thing: catching it lying. A claim more than five seconds over the token's
actual age is flagged `dwell_implausible`. (This rule caught the project's own
smoke test the first time it was written, which is the best evidence available
that it works.)

The **read token** is an HMAC over `{chapter id, story id, required dwell,
issue time, address hash}`, signed with `VIEW_SALT`. It cannot be minted,
edited or backdated by a client. It is bound to the reader's address hash so
that one fetch cannot be fanned out across a botnet — each address has to fetch
the chapter for itself, which makes fetch volume and heartbeat volume
comparable in the logs. A token older than six hours is a replay, not a slow
reader.

### 2. Deduplication: one reader, one chapter, once per 24 hours

- **Signed in:** dedupe on the member id.
- **Signed out:** dedupe on `sha256(VIEW_SALT + ip + "|" + user_agent)`.

The raw IP address is **never stored**, anywhere, in any table. Only the hash
is, and it is useless outside this site and meaningless once the salt is
rotated. Including the user agent means two people behind one household NAT are
usually two readers, which is the honest outcome.

There are two enforcement layers, on purpose:

1. A **rolling 24-hour check** — the actual rule. Stricter than a calendar day,
   because it also catches 23:59 followed by 00:01.
2. A **unique index** on `(dedupe_key, chapter_id, view_day)` — the race
   backstop. The rolling check is a read-then-write, and two heartbeats
   arriving in the same millisecond both pass it. The index cannot be raced.
   When the index wins, the event row is corrected to `duplicate_race` so the
   audit trail says what actually happened.

### 3. Rate caps, counted in Postgres

| Cap | Default | Why that number |
|---|---|---|
| per member, per hour | 60 qualified views | A person reading 40 chapters in an hour is already at the edge of plausible. 60 leaves room for a genuine binge. |
| per address hash, per hour | 40 | Households share addresses; farms share them harder. |
| per story, per hour | 500 | A real hit will reach this. That is intended. |

The per-address cap uses `sha256(VIEW_SALT + ip)` — **without** the user agent.
That is deliberate and it matters: rotating the user agent is the first thing a
farm does, and it must not buy a fresh rate budget. (Dedupe includes the UA;
rate capping excludes it. Two hashes, two jobs.)

These are counted with indexed queries against `qualified_views`, **not** in
an in-memory bucket. This is money; it has to survive a container restart. The
in-memory limiter in `lib/http.js` is for password guessing and write floods
and is a different mechanism entirely.

When the per-story cap bites, the overflow is not silently discarded — it is
recorded raw and flagged, and an admin decides. Losing an hour of a real hit
costs one author some money. Paying out an hour of a botnet costs everyone the
ad account.

### 4. Fatal flags — any one of these disqualifies

Each means "this cannot be a person reading this chapter for the first time
today".

| Flag | Meaning |
|---|---|
| `token_invalid` | no token, forged, or issued for a different chapter |
| `token_stale` | issued longer ago than a reading session can last |
| `token_ip_mismatch` | fetched by one address, heartbeat from another |
| `dwell_too_short` | arrived before the chapter could have been read |
| `dwell_implausible` | claims a longer dwell than the token has existed |
| `duplicate_24h` | this identity already counted for this chapter today |
| `self_view` | the author reading their own story |
| `story_not_payable` | not published, or the author is suspended |
| `no_scroll` | scroll depth below 50% of the chapter |
| `no_interaction` | not one scroll, key, pointer or visibility event |
| `bot_user_agent` | says it is a bot, or is a well-known HTTP client |
| `private_ip` | impossible from the public internet through Caddy |
| `rate_member` / `rate_ip` / `rate_story` | over an hourly cap |
| `uniform_timing` | interaction gaps too regular to be a person |

### 5. Advisory flags — two together disqualify

This is the part that keeps the system from being a blunt instrument.

| Flag | Why it is not fatal alone |
|---|---|
| `datacenter_ip` | a corporate VPN egresses from AWS all day |
| `missing_referrer` | privacy extensions strip it; some apps never send it |
| `foreign_referrer` | not our pages and not a search engine or social network anyone has heard of — where traffic exchanges and pop-unders live |
| `thin_interaction` | scrolled, but barely touched |
| `no_user_agent` | unusual, not impossible |

**A datacenter address alone is a corporate VPN. A missing referrer alone is a
privacy extension. A datacenter address *with* no referrer is a bot** — and no
single rule would have caught it. That combination rule is where most real
automated traffic actually dies.

### 6. Interaction timing

The client reports the gaps between the reader's **interaction events** —
scrolls, keypresses, pointer events, tab focus changes — and the server takes
their coefficient of variation (standard deviation ÷ mean).

It is emphatically **not** the gap between timer ticks. An earlier draft
measured heartbeat timer intervals and flagged every honest reader in the test
suite, because a `setInterval` fires just as regularly for a person as for a
script. The gaps between real human interactions are wildly irregular — a
paragraph goes by in two seconds, then twenty seconds of thinking — and a
generator emitting events on a schedule is not.

Below cv 0.05 over at least 4 samples is `uniform_timing`. Fewer than 4 samples
is not held against the reader.

This is the **easiest rule in the document to defeat** — adding jitter is one
line for an attacker. It is here because it catches the lazy majority cheaply,
not because it is hard to beat.

### 7. Datacenter detection is cheap and knows it

`server/lib/netclass.js` holds a static list of about 120 published cloud
prefixes (AWS, GCP, Azure, DigitalOcean, Hetzner, OVH — including this box's own
provider — Linode, Vultr, Oracle, Contabo, Scaleway, Alibaba, Tencent), checked
with integer arithmetic. No lookup, no API, no latency, no cost, and no reader
IP addresses sent to a third party.

**What it is not:** an IP intelligence feed. A small VPS provider, a residential
proxy network, or a consumer VPN reads as `unknown` and passes. The label
`residential` means only "a public address not on our list" — nothing was
verified. This is stated in the code because treating the list as authoritative
would be worse than not having it: it would produce confident false negatives.

Upgrade path when the money justifies it: pull the providers' published prefix
JSON on a cron into a table and query that instead. Same interface.

IPv6 is classified `unknown`. A static list of IPv6 provider prefixes honest
enough to act on does not fit in this file.

### 8. Manual review threshold

Automated detection catches the obvious and misses the patient. So the amount
of money that can leave without a human looking at it is capped.

A payout line is `held_for_review` — regardless of everything above — when any
of these is true:

| Trigger | Default | Reasoning |
|---|---|---|
| gross ≥ threshold | $50 per period | High enough that a modestly successful author is paid without ceremony; low enough that a farm cannot drain the account one period at a time while staying quiet. |
| rejection rate > 35% of the author's read events | | Not proof of anything — it can be one bad referrer source — but exactly the case where the automated verdict should not be the last word. |
| qualified views > 8× the author's previous best period | | Genuine growth exists. A 20× month does not arrive quietly, and the cost of asking is one message. |
| first period and > 5,000 qualified views | | Everybody starts somewhere. A debut with 40,000 views is not a debut. |
| the member has a `review_note` set | | An author already under review is always held. |

Held lines sit in the admin queue with the reason written out in words, for
whoever opens it next week.

### 8a. Imported works are counted, and are never paid for

The catalogue holds two kinds of work. Novels by authors the owner contracted,
and public-domain or open-licensed books loaded by `server/import.mjs` — see
[`content-sources.md`](content-sources.md).

Both are read. Both carry advertising. Only the first is owed a per-view
payment: nobody is owed a royalty on a novel whose author died in 1817, and
paying the importing account for it would divert the pool away from the people
it is for.

The rule is enforced in **one place, and it is the place where a view becomes
money** — the sweep in `runPeriod`:

```sql
UPDATE qualified_views qv SET payout_period_id = $1
 WHERE qv.payout_period_id IS NULL
   AND qv.created_at >= $2 AND qv.created_at < $3
   AND NOT EXISTS (SELECT 1 FROM story_provenance sp WHERE sp.story_id = qv.story_id)
```

Three deliberate consequences.

**There is no `payout_eligible` boolean anywhere.** A boolean is a second copy
of the truth, and second copies drift; the existence of the
`story_provenance` row *is* the answer. That row also holds the attribution the
licence requires, so making an imported work payable would mean deleting the
credit — which is not something anyone does by accident.

**Qualification itself is untouched.** An imported work's views still go
through every rule above and are still banked in `qualified_views`. That is on
purpose: invalid traffic against a public-domain novel is still invalid traffic,
and AdSense does not care who wrote the book. Excluding these reads from the
fraud aggregation would blind the site to a bot farm pointed at the free half
of the catalogue.

**Those rows are simply never claimed.** They keep a NULL `payout_period_id`
for ever, and stay visible in every count, every stat page and every fraud
query.

The author-side numbers behind the review decision (`raw_views`,
`flagged_events`) are computed over the author's *payable* stories only, so the
rejection ratio — the single most useful fraud signal there is — stays a ratio
of the same population as the qualified count.

### 9. Both counts are stored, always

`stories.raw_view_count` and `stories.qualified_view_count`; `read_events` and
`qualified_views`; `payout_lines.raw_views`, `.qualified_views` and
`.flagged_events`.

The counts are recorded on the payout line **at the time the period ran**,
because recomputing them later gives a different answer once the salt has been
rotated or old events have been pruned.

Authors see both numbers, side by side, always — never only the qualified
figure. An author whose raw count is ten times their qualified count needs to
see that. Either somebody is sending them junk traffic they did not ask for, or
they know exactly where it came from. Hiding the gap only delays that
conversation.

---

## What this does not catch

Stated plainly, because a fraud system trusted past its evidence is worse than
one nobody trusts.

**A patient, well-funded attacker will get through.** Residential proxies,
real browser automation, jittered timing, faked scroll events, one view per
address per chapter per day, spread across enough addresses. Every individual
rule above passes. `server/test/views.test.mjs` contains a test that asserts
exactly this and says so in its comment, so nobody discovers it by surprise.

What stands between that attacker and the money is not `qualify()`. It is:

1. the **per-address hourly cap**, which forces them to buy proxy breadth
   rather than depth;
2. the **payout review threshold**, which caps unattended loss per period;
3. the **raw-versus-qualified ratio** an admin sees in the fraud queue, which
   is very hard to make look normal at scale;
4. the fact that anything at volume shows up in `read_events` as a shape — one
   address hash across many stories, or many address hashes with identical
   behaviour — and the admin fraud endpoint groups by exactly those.

**Other known gaps:**

- **IPv6 readers get no datacenter classification.** They can still be caught
  by every other rule.
- **Shared NAT under-counts.** A school or an office behind one address shares
  a rate cap. The user-agent component of the dedupe hash softens this for
  deduplication but not for the cap.
- **Ad-blocking readers still count as views but generate no ad revenue.** The
  payout rate is set per period from what the ads actually earned, which is
  where this is absorbed rather than being modelled per view.
- **Nothing here is a defence against an author paying real people to read.**
  That is indistinguishable from an author being popular, and arguably it is
  the same thing.

---

## Tuning

Every threshold is an environment variable, so a change is a config edit and a
restart, not a deploy. Defaults are the values in the tables above.

```
READ_WPM_CEILING            600
MIN_DWELL_MS                15000
MAX_DWELL_MS                180000
MAX_SESSION_MS              21600000     # 6h
DEDUPE_MS                   86400000     # 24h
CAP_MEMBER_PER_HOUR         60
CAP_IP_PER_HOUR             40
CAP_STORY_PER_HOUR          500
MIN_SCROLL_DEPTH            0.5
MIN_INTERACTIONS            2
MIN_TIMING_CV               0.05
MIN_TIMING_SAMPLES          4
MAX_ADVISORY_FLAGS          2

PAYOUT_RATE_PER_1K_MICROS         1500000    # $1.50 per 1000 qualified views
PAYOUT_REVIEW_THRESHOLD_MICROS    50000000   # $50
PAYOUT_MAX_FLAG_RATIO             0.35
PAYOUT_MAX_GROWTH_MULTIPLE        8
PAYOUT_NEW_AUTHOR_VIEW_CEILING    5000
```

**`VIEW_SALT` is not a threshold and rotating it is not free.** Rotating it
invalidates every stored visitor hash, which means every returning signed-out
reader looks new for 24 hours and can be counted a second time on a chapter
they already read. Rotate it deliberately, at a period boundary, and expect a
one-day bump in the numbers.

## Where to look when something is wrong

```bash
# what kinds of junk are arriving, where it is landing, and from how many
# distinct addresses — the three views a human needs before approving a payout
curl -H "X-Admin-Token: $ADMIN_TOKEN" \
  'https://stories.cstsolution.com/api/v1/admin/fraud?hours=168'
```

The response has three sections and they are only meaningful together: the flag
histogram says what kind of junk it is, `byStory` says whose story it is
landing on, and `byAddress` says whether it is one machine or a thousand. Any
one alone is easy to explain away.
