Skip to content
SIH Buddyby Ganeev Singh
Dev

๐Ÿ”ฅ Roast My Pick ยท SIH26079

AI-Based Forecast Bust Detection for Medium-Range Weather Forecasts

Ministry of Earth Sciences (MoES)

Mild17/100

Good pick. Genuinely. Now sit down, because the judges are going to try anyway โ€” and this is what they will try.

Strong pick. A genuinely clean supervised problem with free real data, an objectively computable target and a fair baseline โ€” rare on this portal โ€” so define your bust threshold honestly and make beating ensemble spread the whole claim. Roughly 140โ€“330 teams are expected to go here.

The receipts

Every red flag on this statement, in full. These are the four places it bites.

  1. Exhibit A

    Ensemble spread already encodes much of the predictable uncertainty and is a stronger baseline than it looks, so a model that merely correlates with spread has added nothing and you must show the increment

  2. It gets worse

    Busts are by definition rare, so the class imbalance is severe and a model optimised on accuracy will confidently declare every forecast reliable

  3. Still reading?

    Bust is never defined in the statement, so you set the error threshold that determines your own positive class โ€” pick it from operational impact rather than from what makes your numbers look good, and say which you did

  4. And the finisher

    Some forecast failures are genuine limits of atmospheric predictability rather than detectable model weaknesses, so there is a ceiling on this task and claiming to have beaten it should make you suspicious of your own evaluation

The damage report

Every score this statement earned, and what each one actually costs you.

  • Feasibility

    4/5

    Actually buildable, which on this slate is rarer than it sounds. Do not squander it on scope.

    The data position is genuinely good โ€” global forecast archives are openly distributed, research ensemble archives are freely available, and reanalysis provides the verifying truth, so you can construct a real history of forecast errors and learn from it rather than simulating anything.

  • Innovation scope

    4/5

    There is something genuinely new here. Do not bury it under another dashboard.

    The statement names the outputs and nothing about method, so how you represent a forecast's state, what features indicate impending failure and how you distinguish genuine predictability limits from model deficiencies are all yours to work out.

  • Clarity

    4/5

    The ask is unambiguous, which quietly removes your favourite excuse.

    Short but well posed โ€” the four expected outputs are enumerated precisely and the framing that you compare current forecast patterns against historical error behaviour is exactly the right formulation, though no accuracy target or definition of how large an error counts as a bust is given.

  • Acceptance potential

    4/5

    Strong footing before you have written a line. Try not to waste it.

    This is a clean, well-posed supervised problem with genuinely free data, an honest hindcast validation against documented busts, and a fair baseline in ensemble spread โ€” and forecast verification is an unglamorous corner that will attract very few teams despite being operationally valuable.

  • Effort

    Heavy

    Heavy. Somebody on this team is not sleeping in week three. Pick who, on purpose.

    Building an aligned forecast-and-verification archive across ten lead times is substantial data engineering, and on top of it sit the error model, the explainability layer, the ensemble spread baseline and a dashboard.

  • Demo-ability

    Medium

    Demoable, if you rehearse it. Nobody rehearses it.

    Flagging a real historical bust before it happened is a genuinely satisfying moment and it is verifiable against the record, but the product is confidence maps and probabilities, which need framing before a general judge appreciates them.

  • Data

    None supplied

    No dataset comes with this one, so every accuracy figure you quote is a number about labels you invented.

    Nothing is provided with the statement. You are sourcing, cleaning and labelling it yourself, and that work is invisible in the demo but very visible in the questions.

The demo they will have already seen

Somewhere around 140โ€“330 teams are heading here, and the description is doing the choosing for most of them. They will read the same brief, reach the same architecture, and build a version of the same demo you are planning. Being correct is the floor. If your five minutes could be swapped with the team before you and nobody in the room would notice, you have not picked badly โ€” you have built predictably, which costs exactly the same and hurts more.

What survives

The ground worth standing on when the questions start.

  • Historical forecast archives and reanalysis are both freely available, so unlike almost everything else in this block you train and validate on real data rather than on a simulator
  • The target variable is objectively computable โ€” you know exactly how wrong every past forecast was โ€” which makes this a genuinely clean supervised problem with no label ambiguity
  • Ensemble spread gives you an operational baseline that is honest and non-trivial, so demonstrating improvement over it is a real result rather than a claim

Nothing here is fatal. It is just the list of places this statement pushes back, and you now get to push there first.

The framing is a joke. The findings are not โ€” they are the same analysis on the statement page, and every line above is attached to a score or a fact in the record. It is one opinion with its reasoning attached, so argue with it before you trust it.