Skip to content
SIH Buddyby Ganeev Singh
Dev
All problem statements
SIH26102Strong pickacceptance 4/5

Development of an AI-powered system to detect anomalies, fraud, and inefficiencies in MPLAD Scheme implementation regd.

MoSPI · Miscellaneous · Software

Genuine public administrative data at national scale with verifiable findings is rare here, and audit reports give you real irregularity patterns to encode — just frame it as outlier detection for review rather than fraud detection, and keep identifiable constituencies out of the demo.

Open dataset ↗

What it actually is

Members of Parliament recommend local development works — roads, community halls, water facilities — funded from a scheme running thousands of projects nationwide through many implementing agencies. With that volume nobody can spot the works that cost far too much, never got built, or were sanctioned twice. The ask is a system that finds those patterns in the scheme's own data.

What to build

An anomaly detection layer over the scheme's published works data, framed honestly as outlier detection rather than fraud detection since no confirmed fraud labels exist: cost-per-asset outlier detection comparing similar works within and across districts so a community hall costing several times the district norm surfaces; near-duplicate work detection catching the same work description sanctioned at the same or adjacent locations more than once; execution stall detection finding works sanctioned and funded long ago with no progress recorded; utilisation pattern analysis flagging agencies or districts whose expenditure and completion behaviour departs from peers; and a reviewer interface where every flag is explained in terms of the specific comparison that produced it, with the pattern library grounded in the irregularity types that public audit reports on this scheme have already documented.

Smallest thing that wins the room

Show the ten highest-scoring flagged works with the reason beside each — this asset costs four times the district median, this description appears twice at the same location, this work has been funded and stalled for three years — and let a judge click into one and see the comparison set that produced it.

How crowded this one gets

A guess, projected from the 2025 statements — the last year where both the submission counts and the winners were published.

Moderate160–360 teams expectedroughly 1 in 132–305 wins it

Quieter than 24% of the 226 · #173 of 226 by expected field

A normal-sized field. Your idea has to be good, not miraculous.

Why: central ministry statements sat below the average.

This is a guess, not a fact

Nobody has published 2026’s numbers yet. This is an analysed estimate from last year’s pattern, so please do not take it as the truth — check the live counter on the SIH portal before you decide anything. The range covers the middle half of likely outcomes, so one statement in two lands outside it. Entry closes at 500 ideas per statement, so no range goes past that — a statement that reaches the cap fills and shuts rather than drawing an unlimited crowd. The model reads only three things a team can see before choosing — software or hardware, the theme, and what kind of body posted it — and those explain about a quarter of the variation in last year’s field sizes (R² 0.25 on held-out statements). Trust the band more than the number, and the ordering more than either. It cannot see how good your idea is, which is the part that actually decides it.

The scores

The number is the shorthand. The line under it is the reason.

What you will be writing

  • cost-per-unit outlier detection within peer groups
  • near-duplicate work description matching
  • isolation forest on utilisation patterns
  • stalled execution detection from progress timelines
  • explainable flag attribution to comparison set
  • public MPLADS works data ingestion
  • Public expenditure oversight
  • Anomaly detection
  • Development scheme monitoring

Prior art to read before you start

administrative data anomaly detection · duplicate work and cost outlier identification · explainable risk flagging for auditors

Analysed by Claude Opus. Every score above is a judgment call with its reasoning attached — kindly cross-check this against the official statement on the SIH portal before your team commits to it.