Skip to content
SIH Buddyby Ganeev Singh
Dev

๐Ÿ”ฅ Roast My Pick ยท SIH26102

Development of an AI-powered system to detect anomalies, fraud, and inefficiencies in MPLAD Scheme implementation regd.

MoSPI

Mild12/100

Good pick. Genuinely. Now sit down, because the judges are going to try anyway โ€” and this is what they will try.

Strong pick. Genuine public administrative data at national scale with verifiable findings is rare here, and audit reports give you real irregularity patterns to encode โ€” just frame it as outlier detection for review rather than fraud detection, and keep identifiable constituencies out of the demo. Roughly 160โ€“360 teams are expected to go here.

The receipts

Every red flag on this statement, in full. These are the four places it bites.

  1. Exhibit A

    There are no confirmed fraud labels, so this is unsupervised outlier detection and calling it fraud detection is a claim you cannot support โ€” a flagged work is a work worth looking at, nothing more, and framing it otherwise is both wrong and unfair to the people named

  2. It gets worse

    Legitimate cost variation is enormous โ€” terrain, materials, remoteness and timing all move unit costs โ€” so comparison groups have to be constructed carefully or you will flag every hill district

  3. Still reading?

    Naming specific constituencies or members in a demo attaches an allegation to real identifiable people on the basis of a statistical outlier, so anonymise the presentation and let the method be the claim

  4. And the finisher

    The description repeats itself three times and defines nothing, so your detection criteria are your own and a panel may measure you against different ones

The damage report

Every score this statement earned, and what each one actually costs you.

  • Feasibility

    4/5

    Actually buildable, which on this slate is rarer than it sounds. Do not squander it on scope.

    This is one of the few statements on the portal backed by real public administrative data at scale โ€” the scheme portal publishes works, sanctioned amounts, expenditure and completion status across constituencies and districts, so you are analysing genuine government records rather than a synthetic set.

  • Innovation scope

    4/5

    There is something genuinely new here. Do not bury it under another dashboard.

    The statement names no method, no anomaly definition and no model, so what counts as an irregularity, how you construct comparison groups and how you avoid flagging legitimate variation are all yours to determine and are the substance of the work.

  • Clarity

    3/5

    Clear enough to start, vague enough to drift. Write the scope down and stop reinterpreting it weekly.

    The background, description and expected solution sections restate essentially the same paragraph three times, and the statement never defines what an anomaly is or supplies a single confirmed irregular case, so you are setting the detection criteria yourself.

  • Acceptance potential

    4/5

    Strong footing before you have written a line. Try not to waste it.

    Real public administrative data at national scale is rare on this portal and it is the whole game here, the findings are concrete and verifiable against published records, and public audit reports on this scheme document the irregularity patterns you can encode โ€” so your detectors can be grounded in documented reality rather than invented.

  • Effort

    Heavy

    Heavy. Somebody on this team is not sleeping in week three. Pick who, on purpose.

    Acquiring and cleaning the works data, constructing meaningful comparison groups, building several distinct detectors, scoring and ranking flags, and building an explainable reviewer interface is four workstreams with the comparison-group construction being the analytically hardest.

  • Demo-ability

    Medium

    Demoable, if you rehearse it. Nobody rehearses it.

    A ranked list of flagged works with concrete reasons is genuinely compelling to an administrator and each flag is checkable against the source record, but it is a table rather than a visual and needs the comparison logic explained for the ranking to mean anything.

The demo they will have already seen

Somewhere around 160โ€“360 teams are heading here, and the description is doing the choosing for most of them. They will read the same brief, reach the same architecture, and build a version of the same demo you are planning. Being correct is the floor. If your five minutes could be swapped with the team before you and nobody in the room would notice, you have not picked badly โ€” you have built predictably, which costs exactly the same and hurts more.

What survives

The ground worth standing on when the questions start.

  • The data is real, public and at national scale, which puts this in a small minority of statements here where you analyse genuine records rather than something you generated
  • Public audit reports on this scheme document the actual irregularity patterns that have been found historically, so you can encode detectors grounded in what has really gone wrong rather than in what you imagine might
  • Every flag is verifiable โ€” a judge can look up the flagged work in the public portal and see the numbers for themselves, which is a very strong position

Nothing here is fatal. It is just the list of places this statement pushes back, and you now get to push there first.

The framing is a joke. The findings are not โ€” they are the same analysis on the statement page, and every line above is attached to a score or a fact in the record. It is one opinion with its reasoning attached, so argue with it before you trust it.