Program Design

How to pilot a loyalty program before national rollout

A pilot exists to answer one question honestly: if we scale this to 400 districts, what will actually happen? Most pilots are designed — usually by accident — to answer a different question: can we make this look good in our best territory with our best people? This guide covers district selection, sample sizes, control groups, the 8–12 week measurement plan and the pre-agreed decision gates that turn a pilot into a scale decision rather than a debate. It assumes the design work from the 90-day launch playbook is done and the scheme is ready to meet real counters.

Representative or friendly: the district-selection decision

When the pilot is proposed, every sales head instinctively nominates his best district — the one with the energetic ASM, the loyal dealers, the counters that answer the phone. It is an understandable instinct and it quietly ruins the pilot, because every number that comes out is flattered: enrolment is easier, activation is faster, fraud is politer. The national forecast built on that pilot then over-promises by 30–50%, and the program spends year one explaining the gap.

The honest design uses two or three districts in a deliberate mix: one strong (to see the ceiling), one average (to see the likely national reality), and ideally one difficult — weak share, a dominant competitor, patchier field coverage (to see the floor). Match them on the structural variables that drive trade behaviour: urban/rural mix, dealer density, competitor intensity, and category seasonality. A cement pilot in peak monsoon or a lighting pilot in the dead post-Diwali quarter measures the season, not the scheme — check the windows against the festive trade calendar. And pick districts served by different distributors: a single distributor's enthusiasm (or sabotage) is a variable you must not let masquerade as a program result.

Sample sizes and control groups: enough counters to beat the noise

Trade purchase data is noisy — festivals, local projects, copper moves and dealer credit cycles swing counter offtake 20–30% month to month. To read a genuine 8–10% lift through that noise, you need scale: as a working rule, at least 300–500 enrolled counters (or 800–1,000 influencers, whose individual behaviour is lumpier) across the pilot districts. Below that, the pilot review becomes an argument about anecdotes.

The control group is what converts numbers into evidence. Build a matched set of counters — similar pre-pilot purchase volume, outlet type and geography — that stays out of the program, and measure both groups identically through dealer secondary data or field audits. Two practical rules: put controls in an adjacent similar district, not the same town (counters talk; nothing contaminates a control like a neighbour showing off UPI credits), and freeze the matching before launch — selecting controls after you have seen results is the polite name for cheating. The growth gap between enrolled and control counters is your incrementality; the pilot group's own growth alone is marketing. This is the same matched-control method the program will need forever, as covered in the KPI guide — the pilot is where you build the muscle.

One compliance note that pilots forget: Section 194R applies from the first rupee of benefit. Capture PAN at pilot enrolment and run the real TDS machinery — 10% once cumulative benefits cross ₹20,000 per PAN in the financial year — because the pilot is precisely where you want to discover that the TDS pipeline, not just the payout rail, works end to end.

The 8–12 week measurement plan

1

Weeks 0–2: enrolment mechanics

Measure enrolment velocity per field-day, KYC completion rates (where do users abandon — PAN? bank details?), assisted-vs-self enrolment split, and dealer reaction. Target 60–70% of active counters enrolled within four weeks. Every abandonment point found here is friction you fix before multiplying it by 400 districts.

2

Weeks 2–6: activation and habit

First-scan rate within 14 days of enrolment (healthy: 50–70%), scan frequency per active user, payout success rate (>98%) and median scan-to-money time, helpdesk ticket mix, and the first fraud patterns — dealer bulk-scanning shows its face within a fortnight, and the pilot is where the fraud engine's thresholds get tuned against reality rather than assumptions.

3

Weeks 6–12: behaviour change vs control

Now the money question: share-of-wallet movement in enrolled counters vs matched controls, premium-SKU mix shift, repeat-scan retention (are week-2 scanners still scanning in week 10, or was it novelty?), and early redemption behaviour. Two full monthly purchase cycles is the minimum to distinguish habit from launch excitement — which is why pilots shorter than eight live weeks mostly measure enthusiasm.

4

Throughout: unit economics and qualitative truth

Track cost per active counter, reward cost as % of verified secondary, and projected national payback using the ROI calculator. Alongside the dashboard, run structured field interviews in week 4 and week 10 — fifteen counters and ten influencers per district, asked the same questions both times. The dashboard tells you what changed; the counters tell you why, and their exact objections become the national field script.

Decision gates: agree the thresholds before you start

The single highest-leverage act in pilot design is writing the pass/fail thresholds down before launch, signed by the sales head and finance. Typical gates for a counter program:

  • Adoption: 30-day activation above 50%; steady-state monthly actives above 35% of enrolled.
  • Impact: incremental lift vs control of at least 5–8% of secondary volume, with premium-mix movement in the intended direction.
  • Operations: payout success above 98%; median helpdesk resolution under 48 hours; confirmed fraud under 1–2% of reward spend.
  • Economics: pilot unit costs projecting to a national payback of 9–15 months at the budget band in our budgeting guide.

Then the review has only three outcomes: scale (gates passed — roll out in district waves, using later waves as fresh temporary controls), iterate (mechanics work, one gate missed — fix the specific failure and extend the pilot 4–6 weeks; most commonly the reward level or the onboarding flow), or stop (structural failure — the audience does not value the reward, or the channel structure defeats verification). Stopping after a ₹40 lakh pilot is a success of the method: the same failure at national scale costs crores and years of trade goodwill. What pre-agreed gates prevent is the fourth, most common outcome — the ambiguous pilot that scales anyway because momentum and sunk pride demand it.

The pilot traps that flatter results

  • The over-supported pilot. Head office visits weekly, the platform vendor's founder personally unblocks payouts, the ASM does nothing but the pilot. None of this exists in district 247. Ration the support to what the national model will actually provide — one training session, standard helpdesk, normal beat frequency — or run the last four weeks with support deliberately withdrawn and watch what happens.
  • Dealer hand-holding that won't scale. In pilots, dealers often personally walk retailers through scans because the RSM asked them to. At national scale that favour evaporates. Design the pilot so enrolment and scanning survive dealer indifference — because nationally, indifference is the median.
  • The launch-bonus mirage. Extra-generous pilot-only rewards ("double points for pilot districts") produce adoption the national economics cannot repeat. Pilot the real scheme; if you must add launch energy, use time-boxed boosters you can also afford nationally.
  • Measuring the festive quarter. A pilot spanning Diwali will show lift that is seasonality wearing a program badge. Either avoid the window or lean harder on the control group, which experiences the same festival.
  • Cherry-picked reporting. Decide the report template — every gate, both groups, all districts — before launch, and publish the difficult district's numbers with the same prominence as the strong one's.

Frequently asked questions

How many districts should a loyalty pilot cover?

Two or three: one strong territory, one average, and ideally one known to be difficult. A single district cannot separate program effect from local conditions, and more than three multiplies cost and management attention without adding much learning. The mix matters more than the count — three friendly districts produce a beautiful pilot and a misleading national forecast.

How long should a loyalty program pilot run?

Eight to twelve weeks of live scanning, after enrolment stabilises. Shorter than eight weeks and you only measure novelty; you need at least two full monthly purchase cycles to see repeat behaviour, scheme fatigue onset and fraud patterns. Add 4–6 weeks before go-live for serialised stock to reach pilot districts and un-coded inventory to start flushing through.

How big a sample do you need for a credible pilot?

As a working rule, at least 300–500 enrolled counters (or 800–1,000 influencers) across pilot districts, with a matched control set of similar size left out of the program. Below that, normal month-to-month noise in trade purchases — festivals, copper moves, a big local project — swamps the effect you are trying to read, and the pilot ends in argument instead of decision.

What should the control group look like?

Counters matched to enrolled ones on pre-pilot purchase volume, outlet type and geography — ideally in an adjacent, similar district rather than the same town, to avoid contamination when neighbouring counters hear about rewards. Measure both groups identically through dealer secondary data or field audits, and read incrementality as the growth gap between them, not the pilot group's growth alone.

What are the decision gates for scaling a pilot nationally?

Set the thresholds before the pilot starts. Typical gates: 30-day activation above 50%, steady-state monthly actives above 35%, verified incremental lift versus control of at least 5–8% of secondary volume, payout success above 98%, confirmed fraud under 1–2% of reward spend, and unit economics that project to a 9–15 month national payback. Pre-agreed gates turn the scale-up review into a reading, not a debate.

Do pilot rewards attract TDS as well?

Yes — Section 194R does not distinguish pilots from programs. Any participant whose cumulative benefits cross ₹20,000 in the financial year triggers 10% TDS, so capture PAN at pilot enrolment and run the same benefit-aggregation machinery you plan for national scale. The pilot is also the right time to test that the TDS pipeline works end to end.

Pilot in two districts in four weeks

Unotag stands up pilot programs fast — serialised QR for pilot stock, matched-control dashboards, fraud tuning and 194R handling from scan one — with pilot-first commercial terms.

Related reading