๐ฅ Roast My Pick ยท SIH26028
Dynamic Forecast of Expected Time of Arrival (ETA) for Coaching Trains
Ministry of Railways
Good pick. Genuinely. Now sit down, because the judges are going to try anyway โ and this is what they will try.
Strong pick. The best data availability and the most independently verifiable demo of any railway statement here, but the schedule-plus-delay baseline is genuinely competitive, so start collecting running data immediately and make beating that baseline by a stated margin your entire claim. Roughly 140โ330 teams are expected to go here.
The receipts
Every red flag on this statement, in full. These are the four places it bites.
Exhibit A
The baseline is stronger than it looks โ schedule plus current delay plus built-in recovery time is already reasonably good on short horizons, so a model that ties it has achieved nothing and you must report the comparison honestly
It gets worse
Delay propagation is the whole problem and it is cascading and network-wide; a per-train model that ignores the preceding train on the same section will plateau quickly
Still reading?
Collecting a historical running corpus takes calendar time you may not have if you start late, so begin the data collection before you write any modelling code
And the finisher
Multi-day long-distance runs are where errors compound most and are the hardest case, and a team that evaluates only on short suburban runs has demonstrated the easy half
The damage report
Every score this statement earned, and what each one actually costs you.
Feasibility
4/5Actually buildable, which on this slate is rarer than it sounds. Do not squander it on scope.
This is one of the very few railway statements with genuinely obtainable data โ the published timetable, station and route master data and live running status are all accessible, and historical delay records have been collected and studied publicly, so you can train on real journeys and validate against real arrivals rather than a simulation.
Innovation scope
4/5There is something genuinely new here. Do not bury it under another dashboard.
The statement suggests candidate signals but prescribes no architecture at all, so the modelling approach, the feature construction, whether you predict point estimates or distributions, and how you handle cascading multi-day journeys are all yours.
Clarity
4/5The ask is unambiguous, which quietly removes your favourite excuse.
The objective is unambiguous and, unusually, naturally measurable โ arrival time error in minutes against a stated baseline โ and the delivery requirements including APIs and network-wide scale are explicit, though no accuracy target is named.
Acceptance potential
4/5Strong footing before you have written a line. Try not to waste it.
Real obtainable data, an objectively measurable target, a demo a judge can independently verify and obvious daily-life impact make this genuinely strong, with the only real drag being that train delay prediction has been attempted before and you must beat the naive baseline convincingly rather than merely build a model.
Effort
HeavyHeavy. Somebody on this team is not sleeping in week three. Pick who, on purpose.
A running-status collection pipeline, historical dataset assembly, feature engineering over sectional behaviour, model training and evaluation, a continuously updating serving layer and APIs is five workstreams, and assembling a clean historical delay corpus is slower than the modelling.
Demo-ability
EasyEasy to demo โ and so is everyone else's. Working is the floor here, not the achievement.
This is the rare demo a judge can verify independently โ they can check your prediction against the actual arrival on their own phone โ and being right in front of a sceptical audience is far more persuasive than any interface.
Data
None suppliedNo dataset comes with this one, so every accuracy figure you quote is a number about labels you invented.
Nothing is provided with the statement. You are sourcing, cleaning and labelling it yourself, and that work is invisible in the demo but very visible in the questions.
The demo they will have already seen
Somewhere around 140โ330 teams are heading here, and the description is doing the choosing for most of them. They will read the same brief, reach the same architecture, and build a version of the same demo you are planning. Being correct is the floor. If your five minutes could be swapped with the team before you and nobody in the room would notice, you have not picked badly โ you have built predictably, which costs exactly the same and hurts more.
What survives
The ground worth standing on when the questions start.
- The success metric is naturally quantified in minutes of error, so unlike most statements you can prove you succeeded instead of arguing it
- A judge can independently check a live prediction against the real arrival, which is an extraordinarily strong demo property that almost nothing else on the portal offers
- Historical Indian Railways running data has been collected and analysed publicly, so you can train on real journeys instead of a simulator
Nothing here is fatal. It is just the list of places this statement pushes back, and you now get to push there first.
The framing is a joke. The findings are not โ they are the same analysis on the statement page, and every line above is attached to a score or a fact in the record. It is one opinion with its reasoning attached, so argue with it before you trust it.