๐ฅ Roast My Pick ยท SIH26069
National Weather Big Data Analytics Platform
Ministry of Earth Sciences (MoES)
Reasonable choice. The scoreboard liked it. The scoreboard is not the one asking questions on the day.
Worth considering. The verification idea is genuinely good and cross-checking citizen claims against real meteorological data is the differentiator โ but sort out where your reports will actually come from before you write any code, because the statement assumes access that no longer exists. Roughly 140โ330 teams are expected to go here.
The receipts
Every red flag on this statement, in full. These are the four places it bites.
Exhibit A
Bulk access to the major social platforms is now paid or restricted and scraping breaches their terms, so the primary data source is the weakest part of the plan โ line up an alternative such as public citizen-report feeds or an open platform before you commit
It gets worse
Fake report detection is stated in one line and is by far the hardest requirement here; a team that builds ingestion and a dashboard has skipped the only part of this statement that is not commodity engineering
Still reading?
Absence of a rainfall signal is not proof a report is false โ highly localised events genuinely do fall between gauge and grid resolution, so your system should downgrade confidence rather than declare a report fake
And the finisher
Geotagged metadata is frequently stripped or falsified on social platforms, so location often has to be inferred from text and imagery, which is a separate and error-prone problem the statement does not acknowledge
The damage report
Every score this statement earned, and what each one actually costs you.
Feasibility
3/5Buildable. Not comfortably. There is a week in here you have not planned for yet.
The processing side is straightforward and meteorological data for cross-checking claims is genuinely public, but the collection side is the problem โ bulk programmatic access to the major social platforms is now either expensive or restricted and scraping them breaches their terms, so the primary source the statement is built around is not freely available.
Innovation scope
3/5Mildly interesting. The novelty will not carry the room; the build has to.
The pipeline, the metadata, the machine learning tasks and the dashboard filters are all named, but how you actually verify a report โ which is the hard part and the point of the system โ is left completely open.
Clarity
3/5Clear enough to start, vague enough to drift. Write the scope down and stop reinterpreting it weekly.
The ingestion, the event categories and the dashboard features are described clearly, but the verification requirement sits in a single line despite being far harder than everything else combined, and no accuracy target or definition of a verified report is given anywhere.
Acceptance potential
3/5Middle of the pack. This statement will not win the room for you โ you will have to.
Crowdsourced disaster reporting is a familiar shape, but the verification angle rescues it โ cross-checking a citizen claim against the actual meteorological record for that place and time is a genuinely clever and demonstrable idea that most teams building this will skip in favour of a nicer map.
Effort
HeavyHeavy. Somebody on this team is not sleeping in week three. Pick who, on purpose.
Multi-source ingestion, media and text deduplication, event classification, cross-referenced plausibility checking and a filtered analytics dashboard is five components, with the verification layer alone being a substantial piece of work.
Demo-ability
MediumDemoable, if you rehearse it. Nobody rehearses it.
A map filling with citizen reports is engaging and the moment a recycled photograph gets caught is genuinely satisfying, but the demo depends entirely on you having curated a batch that contains the failure cases worth catching.
Data
None suppliedNo dataset comes with this one, so every accuracy figure you quote is a number about labels you invented.
Nothing is provided with the statement. You are sourcing, cleaning and labelling it yourself, and that work is invisible in the demo but very visible in the questions.
The demo they will have already seen
Somewhere around 140โ330 teams are heading here, and the description is doing the choosing for most of them. They will read the same brief, reach the same architecture, and build a version of the same demo you are planning. Being correct is the floor. If your five minutes could be swapped with the team before you and nobody in the room would notice, you have not picked badly โ you have built predictably, which costs exactly the same and hurts more.
What survives
The ground worth standing on when the questions start.
- Cross-referencing a citizen claim against the independent meteorological record for that location and time is the strongest verification signal available and it is entirely buildable from public rainfall and satellite data โ this is the idea that makes the project worth doing
- Perceptual hashing catches recycled disaster photographs reliably and cheaply, and recirculated imagery from past events is genuinely the dominant form of weather misinformation
- The event categories are enumerated in the statement, so your classification taxonomy is fixed by the sponsor rather than chosen
Nothing here is fatal. It is just the list of places this statement pushes back, and you now get to push there first.
The framing is a joke. The findings are not โ they are the same analysis on the statement page, and every line above is attached to a score or a fact in the record. It is one opinion with its reasoning attached, so argue with it before you trust it.