Skip to content
SIH Buddyby Ganeev Singh
Dev

πŸ”₯ Roast My Pick Β· SIH26013

Automated lntegration and lntelligent Harmonization of Multi-source Geospatial Data for urban Land Record Management.

Ministry of Rural Development

Medium56/100

Bold. Let us find out precisely how bold, in the order a panel will find out.

Proceed with caution. A genuinely important and uncrowded problem, but you will be grading your own homework on synthetic conflicts and the result is invisible infrastructure β€” only take it if you can stage the before-and-after so a non-GIS judge feels the difference. Roughly 140–330 teams are expected to go here.

The receipts

Every red flag on this statement, in full. These are the four places it bites.

  1. Exhibit A

    You cannot obtain genuinely conflicting departmental datasets, so you inject the misalignments and then resolve them β€” validating against your own synthetic conflicts proves your code runs, not that it works on real messy data

  2. It gets worse

    The description lists ten input dataset types and each format you support is real parsing work, so a team that tries to cover all ten ships ten shallow readers

  3. Still reading?

    There is no working artifact a non-specialist judge can appreciate; without a very carefully staged before-and-after this reads as backend plumbing

  4. And the finisher

    Harmonised is never defined, so you are simultaneously setting the requirement and claiming to have met it, which is a weak position under questioning

The damage report

Every score this statement earned, and what each one actually costs you.

  • Feasibility

    3/5

    Buildable. Not comfortably. There is a week in here you have not planned for yet.

    The processing stack is all open β€” GDAL, PostGIS, GeoPandas and standard spatial join logic β€” but real conflicting departmental datasets for an Indian city are not obtainable, so you construct the misalignments yourself and then demonstrate that your engine resolves the misalignments you created.

  • Innovation scope

    4/5

    There is something genuinely new here. Do not bury it under another dashboard.

    Spatial entity matching, conflict resolution and confidence scoring are genuinely open research problems and the description names capabilities without prescribing any method, so the matching algorithm and the confidence model are entirely your design.

  • Clarity

    3/5

    Clear enough to start, vague enough to drift. Write the scope down and stop reinterpreting it weekly.

    The input datasets and the capability list are both explicit, but there is no definition of what harmonised means as an outcome, no acceptance criteria for a correct match and no specification of what confidence scoring measures, so the success condition is left entirely to you.

  • Acceptance potential

    3/5

    Middle of the pack. This statement will not win the room for you β€” you will have to.

    The problem is real and unglamorous enough to thin the field, but you are validating your reconciliation engine against conflicts you manufactured, which is the circular setup that a sharp judge will identify, and the output is infrastructure rather than something anyone can see working.

  • Effort

    Heavy

    Heavy. Somebody on this team is not sleeping in week three. Pick who, on purpose.

    Ten input formats, a coordinate transformation engine, spatial matching, topology correction, attribute schema mapping, change detection, conflict resolution and confidence scoring is eight components, and each new input format quietly costs a day of parsing work.

  • Demo-ability

    Medium

    Demoable, if you rehearse it. Nobody rehearses it.

    Data harmonisation is intrinsically unphotogenic β€” a before-and-after overlay of misaligned polygons snapping into agreement reads well, but only to a judge who already understands why the misalignment was a problem.

  • Data

    None supplied

    No dataset comes with this one, so every accuracy figure you quote is a number about labels you invented.

    Nothing is provided with the statement. You are sourcing, cleaning and labelling it yourself, and that work is invisible in the demo but very visible in the questions.

The demo they will have already seen

Somewhere around 140–330 teams are heading here, and the description is doing the choosing for most of them. They will read the same brief, reach the same architecture, and build a version of the same demo you are planning. Being correct is the floor. If your five minutes could be swapped with the team before you and nobody in the room would notice, you have not picked badly β€” you have built predictably, which costs exactly the same and hurts more.

What survives

The ground worth standing on when the questions start.

  • Data engineering statements attract far fewer teams than model-building ones, so the competitive field here is thin
  • The NAKSHA programme framing means the sponsor has a live operational need rather than a hypothetical one, and a DoLR judge will recognise the workflow immediately
  • The confidence scoring requirement is an unusual and genuinely defensible differentiator β€” most teams will silently auto-resolve conflicts and yours can show its uncertainty

Nothing here is fatal. It is just the list of places this statement pushes back, and you now get to push there first.

The framing is a joke. The findings are not β€” they are the same analysis on the statement page, and every line above is attached to a score or a fact in the record. It is one opinion with its reasoning attached, so argue with it before you trust it.