Skip to content
SIH Buddyby Ganeev Singh
Dev

๐Ÿ”ฅ Roast My Pick ยท SIH26069

National Weather Big Data Analytics Platform

Ministry of Earth Sciences (MoES)

Mild33/100

Reasonable choice. The scoreboard liked it. The scoreboard is not the one asking questions on the day.

Worth considering. The verification idea is genuinely good and cross-checking citizen claims against real meteorological data is the differentiator โ€” but sort out where your reports will actually come from before you write any code, because the statement assumes access that no longer exists. Roughly 140โ€“330 teams are expected to go here.

The receipts

Every red flag on this statement, in full. These are the four places it bites.

  1. Exhibit A

    Bulk access to the major social platforms is now paid or restricted and scraping breaches their terms, so the primary data source is the weakest part of the plan โ€” line up an alternative such as public citizen-report feeds or an open platform before you commit

  2. It gets worse

    Fake report detection is stated in one line and is by far the hardest requirement here; a team that builds ingestion and a dashboard has skipped the only part of this statement that is not commodity engineering

  3. Still reading?

    Absence of a rainfall signal is not proof a report is false โ€” highly localised events genuinely do fall between gauge and grid resolution, so your system should downgrade confidence rather than declare a report fake

  4. And the finisher

    Geotagged metadata is frequently stripped or falsified on social platforms, so location often has to be inferred from text and imagery, which is a separate and error-prone problem the statement does not acknowledge

The damage report

Every score this statement earned, and what each one actually costs you.

  • Feasibility

    3/5

    Buildable. Not comfortably. There is a week in here you have not planned for yet.

    The processing side is straightforward and meteorological data for cross-checking claims is genuinely public, but the collection side is the problem โ€” bulk programmatic access to the major social platforms is now either expensive or restricted and scraping them breaches their terms, so the primary source the statement is built around is not freely available.

  • Innovation scope

    3/5

    Mildly interesting. The novelty will not carry the room; the build has to.

    The pipeline, the metadata, the machine learning tasks and the dashboard filters are all named, but how you actually verify a report โ€” which is the hard part and the point of the system โ€” is left completely open.

  • Clarity

    3/5

    Clear enough to start, vague enough to drift. Write the scope down and stop reinterpreting it weekly.

    The ingestion, the event categories and the dashboard features are described clearly, but the verification requirement sits in a single line despite being far harder than everything else combined, and no accuracy target or definition of a verified report is given anywhere.

  • Acceptance potential

    3/5

    Middle of the pack. This statement will not win the room for you โ€” you will have to.

    Crowdsourced disaster reporting is a familiar shape, but the verification angle rescues it โ€” cross-checking a citizen claim against the actual meteorological record for that place and time is a genuinely clever and demonstrable idea that most teams building this will skip in favour of a nicer map.

  • Effort

    Heavy

    Heavy. Somebody on this team is not sleeping in week three. Pick who, on purpose.

    Multi-source ingestion, media and text deduplication, event classification, cross-referenced plausibility checking and a filtered analytics dashboard is five components, with the verification layer alone being a substantial piece of work.

  • Demo-ability

    Medium

    Demoable, if you rehearse it. Nobody rehearses it.

    A map filling with citizen reports is engaging and the moment a recycled photograph gets caught is genuinely satisfying, but the demo depends entirely on you having curated a batch that contains the failure cases worth catching.

  • Data

    None supplied

    No dataset comes with this one, so every accuracy figure you quote is a number about labels you invented.

    Nothing is provided with the statement. You are sourcing, cleaning and labelling it yourself, and that work is invisible in the demo but very visible in the questions.

The demo they will have already seen

Somewhere around 140โ€“330 teams are heading here, and the description is doing the choosing for most of them. They will read the same brief, reach the same architecture, and build a version of the same demo you are planning. Being correct is the floor. If your five minutes could be swapped with the team before you and nobody in the room would notice, you have not picked badly โ€” you have built predictably, which costs exactly the same and hurts more.

What survives

The ground worth standing on when the questions start.

  • Cross-referencing a citizen claim against the independent meteorological record for that location and time is the strongest verification signal available and it is entirely buildable from public rainfall and satellite data โ€” this is the idea that makes the project worth doing
  • Perceptual hashing catches recycled disaster photographs reliably and cheaply, and recirculated imagery from past events is genuinely the dominant form of weather misinformation
  • The event categories are enumerated in the statement, so your classification taxonomy is fixed by the sponsor rather than chosen

Nothing here is fatal. It is just the list of places this statement pushes back, and you now get to push there first.

The framing is a joke. The findings are not โ€” they are the same analysis on the statement page, and every line above is attached to a score or a fact in the record. It is one opinion with its reasoning attached, so argue with it before you trust it.