Skip to content
SIH Buddyby Ganeev Singh
Dev

๐Ÿ”ฅ Roast My Pick ยท SIH26028

Dynamic Forecast of Expected Time of Arrival (ETA) for Coaching Trains

Ministry of Railways

Mild17/100

Good pick. Genuinely. Now sit down, because the judges are going to try anyway โ€” and this is what they will try.

Strong pick. The best data availability and the most independently verifiable demo of any railway statement here, but the schedule-plus-delay baseline is genuinely competitive, so start collecting running data immediately and make beating that baseline by a stated margin your entire claim. Roughly 140โ€“330 teams are expected to go here.

The receipts

Every red flag on this statement, in full. These are the four places it bites.

  1. Exhibit A

    The baseline is stronger than it looks โ€” schedule plus current delay plus built-in recovery time is already reasonably good on short horizons, so a model that ties it has achieved nothing and you must report the comparison honestly

  2. It gets worse

    Delay propagation is the whole problem and it is cascading and network-wide; a per-train model that ignores the preceding train on the same section will plateau quickly

  3. Still reading?

    Collecting a historical running corpus takes calendar time you may not have if you start late, so begin the data collection before you write any modelling code

  4. And the finisher

    Multi-day long-distance runs are where errors compound most and are the hardest case, and a team that evaluates only on short suburban runs has demonstrated the easy half

The damage report

Every score this statement earned, and what each one actually costs you.

  • Feasibility

    4/5

    Actually buildable, which on this slate is rarer than it sounds. Do not squander it on scope.

    This is one of the very few railway statements with genuinely obtainable data โ€” the published timetable, station and route master data and live running status are all accessible, and historical delay records have been collected and studied publicly, so you can train on real journeys and validate against real arrivals rather than a simulation.

  • Innovation scope

    4/5

    There is something genuinely new here. Do not bury it under another dashboard.

    The statement suggests candidate signals but prescribes no architecture at all, so the modelling approach, the feature construction, whether you predict point estimates or distributions, and how you handle cascading multi-day journeys are all yours.

  • Clarity

    4/5

    The ask is unambiguous, which quietly removes your favourite excuse.

    The objective is unambiguous and, unusually, naturally measurable โ€” arrival time error in minutes against a stated baseline โ€” and the delivery requirements including APIs and network-wide scale are explicit, though no accuracy target is named.

  • Acceptance potential

    4/5

    Strong footing before you have written a line. Try not to waste it.

    Real obtainable data, an objectively measurable target, a demo a judge can independently verify and obvious daily-life impact make this genuinely strong, with the only real drag being that train delay prediction has been attempted before and you must beat the naive baseline convincingly rather than merely build a model.

  • Effort

    Heavy

    Heavy. Somebody on this team is not sleeping in week three. Pick who, on purpose.

    A running-status collection pipeline, historical dataset assembly, feature engineering over sectional behaviour, model training and evaluation, a continuously updating serving layer and APIs is five workstreams, and assembling a clean historical delay corpus is slower than the modelling.

  • Demo-ability

    Easy

    Easy to demo โ€” and so is everyone else's. Working is the floor here, not the achievement.

    This is the rare demo a judge can verify independently โ€” they can check your prediction against the actual arrival on their own phone โ€” and being right in front of a sceptical audience is far more persuasive than any interface.

  • Data

    None supplied

    No dataset comes with this one, so every accuracy figure you quote is a number about labels you invented.

    Nothing is provided with the statement. You are sourcing, cleaning and labelling it yourself, and that work is invisible in the demo but very visible in the questions.

The demo they will have already seen

Somewhere around 140โ€“330 teams are heading here, and the description is doing the choosing for most of them. They will read the same brief, reach the same architecture, and build a version of the same demo you are planning. Being correct is the floor. If your five minutes could be swapped with the team before you and nobody in the room would notice, you have not picked badly โ€” you have built predictably, which costs exactly the same and hurts more.

What survives

The ground worth standing on when the questions start.

  • The success metric is naturally quantified in minutes of error, so unlike most statements you can prove you succeeded instead of arguing it
  • A judge can independently check a live prediction against the real arrival, which is an extraordinarily strong demo property that almost nothing else on the portal offers
  • Historical Indian Railways running data has been collected and analysed publicly, so you can train on real journeys instead of a simulator

Nothing here is fatal. It is just the list of places this statement pushes back, and you now get to push there first.

The framing is a joke. The findings are not โ€” they are the same analysis on the statement page, and every line above is attached to a score or a fact in the record. It is one opinion with its reasoning attached, so argue with it before you trust it.