Skip to content
SIH Buddyby Ganeev Singh
Dev

๐Ÿ”ฅ Roast My Pick ยท SIH26188

Al-Based Fake Identity & Document Screening System

Ministry of Home Affairs

Medium39/100

Reasonable choice. The scoreboard liked it. The scoreboard is not the one asking questions on the day.

Worth considering. Excellently specified with the tampering module as the clear differentiator, but you build your own tampered examples so the detector only catches what you made โ€” invest in the forensic detection, lean on deterministic MRZ validation, and be honest that real forgeries are subtler than synthetic ones. Roughly 150โ€“340 teams are expected to go here.

The receipts

Every red flag on this statement, in full. These are the four places it bites.

  1. Exhibit A

    No public tampered-travel-document dataset exists, so you construct your own manipulations and the detector only catches what you created

  2. It gets worse

    Real forgeries are far subtler than synthetic ones, so a detector trained on your tampering will miss sophisticated real fakes

  3. Still reading?

    Image-forensic cues like error-level analysis are unreliable and easily defeated, so overclaiming tampering detection is risky

  4. And the finisher

    Working with real identity documents raises privacy and data-handling concerns that constrain what you can use

The damage report

Every score this statement earned, and what each one actually costs you.

  • Feasibility

    3/5

    Buildable. Not comfortably. There is a week in here you have not planned for yet.

    OCR, format validation and face verification are standard, and image-forensic tampering detection has established techniques, but the core tampering module is genuinely hard on real forgeries, and there is no public dataset of tampered travel documents, so you construct your own tampered examples and the detector only catches the manipulations you created.

  • Innovation scope

    3/5

    Mildly interesting. The novelty will not carry the room; the build has to.

    OCR, validation and face matching are commodity, so the genuine room is the tampering-detection module the description itself names as the core innovation, which is where the real work and differentiation lie.

  • Clarity

    5/5

    The ask is unambiguous, which quietly removes your favourite excuse.

    The description enumerates the four modules, the document types, the exact fields to extract, and the specific tampering use cases โ€” photo replacement, text manipulation, stamp forgery โ€” making the requirement precisely defined.

  • Acceptance potential

    3/5

    Middle of the pack. This statement will not win the room for you โ€” you will have to.

    The specification is excellent and the tampering focus is the right differentiator, but there is no public tampered-document dataset so you construct your own manipulations and the detector only catches what you made, real forgeries are far subtler than synthetic ones, and an MHA judge will test whether it generalises beyond your examples.

  • Effort

    Massive

    A semester of work wearing a hackathon costume. Something is getting cut; decide what now, not in week five.

    Four modules โ€” OCR, validation, forensic tampering detection and face verification โ€” plus the risk-scoring layer is a broad build, with the tampering detection the deep and demanding part.

  • Demo-ability

    Easy

    Easy to demo โ€” and so is everyone else's. Working is the floor here, not the achievement.

    A genuine document passing and a tampered one flagged with the specific manipulation highlighted and a higher risk score is a concrete, self-explanatory demo.

  • Data

    None supplied

    No dataset comes with this one, so every accuracy figure you quote is a number about labels you invented.

    Nothing is provided with the statement. You are sourcing, cleaning and labelling it yourself, and that work is invisible in the demo but very visible in the questions.

The demo they will have already seen

Somewhere around 150โ€“340 teams are heading here, and the description is doing the choosing for most of them. They will read the same brief, reach the same architecture, and build a version of the same demo you are planning. Being correct is the floor. If your five minutes could be swapped with the team before you and nobody in the room would notice, you have not picked badly โ€” you have built predictably, which costs exactly the same and hurts more.

What survives

The ground worth standing on when the questions start.

  • OCR, MRZ validation and face verification are commodity components with mature libraries, so the surrounding modules are quick
  • The MRZ checksum on passports gives you deterministic validation that catches many alterations without any model
  • The description names tampering detection as the core innovation, so you know exactly where to invest

Nothing here is fatal. It is just the list of places this statement pushes back, and you now get to push there first.

The framing is a joke. The findings are not โ€” they are the same analysis on the statement page, and every line above is attached to a score or a fact in the record. It is one opinion with its reasoning attached, so argue with it before you trust it.