Survey & assessment question sets
+ New question setThe two groups below look alike but do completely different jobs, and they must not be mixed. A survey is how we hear what people say; an assessment measures what changed between the start and the end of the program.
Assessments pay no tokens — enforced at the data layer
A baseline/endline set has no reward field at all. If sitting one paid tokens, people would have a reason to sit it over and over, and the measurement would be worthless. Voluntary surveys pay no tokens either: new tokens come only from completed lessons and fixed Showdown milestones.
Counts refresh overnight.
Group 1 — Voluntary surveys Optional
A survey is offered outside a live run, always with a skip button, and skipping costs nothing. Responses are stored separately and read only in aggregate by the impact analyst.
| Code | Set name | Questions | When it is offered | Tokens | Status | |
|---|---|---|---|---|---|---|
| S-02 | Self-reported behaviour v2.0 real gambling outside the app, stored de-identified |
6 | Month 1 and month 3 | None | In field | View wording |
| S-02 | Self-reported behaviour v2.1 question 4 is double-barrelled — split it in two |
7 | Planned for Q3 2026 | None | Draft · wording with the reviewer | |
| S-01 | Experience of the app v1.0 open questions, none required |
4 | Once, at the end of a period | None | Closed | Read-only |
A survey that asks someone about their own gambling still goes through clinical review, like any other sensitive content. A question such as “do you regret it?” is a judgement rather than a measurement, and it is sent back.
Group 2 — Baseline and endline assessments Versioned
The same set is sat twice: once on joining (baseline) and once after a cycle of the program (endline). The difference between the two is the headline measure of whether tibbi works. The impact analyst reads the results.
Set A-01 · Understanding probability and the house edge
- 20 questions · one answer each · no time limit.
- Correct answers are not shown after submitting — showing them at baseline would teach the person and inflate the endline score.
- No tokens, no badges, no ranking.
- A resit is only possible when the impact analyst reopens the set for a cohort, and resits are marked separately so they never mix into the main figures.
Written here, questions signed off by the clinical reviewer before a cohort ever sees them.
Why versions are locked
If someone sits baseline on v1.2 and endline on v1.3 with different questions, comparing the two means nothing. Every attempt records question_set_version, and endline is forced to use the same version as the baseline.
So a new set only ever applies to a new cohort. There is no “upgrade everyone to the latest version” button.
This is the easiest part of the measurement to break, which is why it is locked rather than left to care.
Versions of set A-01
| Version | Questions | Cohort using it | Attempts | Status | |
|---|---|---|---|---|---|
| v1.2 | 20 | Q2 2026 · 64 people | 105 | In use · locked | View questions |
| v1.3 | 20 | no cohort assigned · planned for Q3 2026 | 0 | Draft · 2 questions reworded | |
| v1.1 | 20 | internal trial, Q1 2026 | 12 | Archived | Read-only |
| v1.0 | 18 | never released | 0 | Archived | Read-only |
v1.2 is locked — editing it directly is blocked
64 people sat baseline on v1.2 and 23 of them have not sat endline yet. Editing v1.2 now would break the comparison for the whole cohort. To change a question, put the rewording into the v1.3 draft and assign that to the next cohort. There is no exception, not even for a typo — a typo is still a change to what the question asks.
What the v1.3 draft changes, line by line
Two questions are reworded and nothing else moves — same twenty questions, same order, same marks. Both changes are here so that when Q3 2026 is reported it is clear exactly what was different about the ruler, without anyone having to open the draft.
| Code | v1.2 · locked wording | v1.3 · draft wording | Reason for the change |
|---|---|---|---|
| B-05 | “In the previous question the total came to more than 100%. What is the extra part called?” | “Two sides of a match are priced so the implied probabilities add up to about 105%. What is the extra 5% called?” | B-05 currently depends on B-04 having been answered first. Anyone who skips B-04 is asked about a total they were never shown, so the question stops measuring the margin and starts measuring memory. |
| B-16 | “Over a long series of bets on a market carrying a margin, what happens to the total staked?” | “Over a long series of bets on a market carrying a margin, what happens to the money you started with?” | Several people in the Q2 cohort read “the total staked” as the running sum of everything they had put on, which rises, rather than the balance, which falls. The wrong answers were about the phrasing, not the maths. |
Neither change goes anywhere near v1.2 — 23 people have not sat endline yet
Both rewordings make the questions easier to read, which also makes them easier to answer. If B-05 and B-16 were changed on v1.2 now, the 23 people still to sit endline would be answering an easier paper than the one they sat at baseline, and the whole Q2 cohort's before-and-after figure would rise for no reason other than the edit. So v1.2 stays frozen word for word until the last endline is in, v1.3 is locked before Q3 opens, and the report states that Q2 and Q3 used different versions and are not pooled into one trend.