๐ฅ Roast My Pick ยท SIH26066
OceanEmbed - Satellite Embedding-Based Deep Learning Framework for Reconstruction of Subsurface Ocean Temperature from Surface Satellite Observations.
Ministry of Earth Sciences (MoES)
Good pick. Genuinely. Now sit down, because the judges are going to try anyway โ and this is what they will try.
Strong pick. The best-posed machine learning statement in this range โ fixed domain, free named datasets and genuinely independent validation โ so spend your first days on the regridding pipeline and report skill by depth honestly rather than hiding the deep levels in an average. Roughly 95โ220 teams are expected to go here.
The receipts
Every red flag on this statement, in full. These are the four places it bites.
Exhibit A
The regridding and harmonisation pipeline is the real project โ six products with different native resolutions, land masks and missing-data conventions will consume more time than training the network, so start there rather than with the model
It gets worse
Reanalysis targets are themselves model output rather than observations, so a model that matches the reanalysis perfectly has learned the reanalysis, and the float validation is the only honest measure of skill
Still reading?
Skill falls off sharply with depth because surface signatures carry less information about the deep ocean, so report by depth level and expect the thousand-metre results to be weak โ a single averaged RMSE hides exactly the failure a judge will ask about
And the finisher
Subsurface reconstruction from satellite data has an established research literature, so be ready to say what your embedding approach adds over the published regression and neural baselines
The damage report
Every score this statement earned, and what each one actually costs you.
Feasibility
5/5Actually buildable, which on this slate is rarer than it sounds. Do not squander it on scope.
Every dataset needed is public and the statement names them โ the reanalysis product for training targets, gridded float data for independent validation, and all six surface satellite variables are freely distributed โ the domain, resolution and depth levels are fixed for you, and the statement even authorises regridding where a product does not match the required resolution.
Innovation scope
3/5Mildly interesting. The novelty will not carry the room; the build has to.
The domain, resolution, input variables, output depths, target and validation datasets and evaluation metrics are all fixed, and even the candidate architecture families are listed, so your latitude sits in the embedding design and how you exploit spatial structure rather than in defining the problem.
Clarity
5/5The ask is unambiguous, which quietly removes your favourite excuse.
One of the most precisely posed machine learning statements on the portal โ it gives geographic bounds, spatial and temporal resolution, the exact input variable list, all fifteen output depth levels in metres, the training target product, the independent validation source and the skill metrics.
Acceptance potential
4/5Strong footing before you have written a line. Try not to waste it.
Everything is defined, all data is free and named, validation is against genuinely independent observations rather than a held-out slice of your own training set, and physical oceanography attracts very few teams โ the caution is that this reconstruction problem has an existing research literature, so you are competing on execution against published baselines rather than on originality.
Effort
HeavyHeavy. Somebody on this team is not sleeping in week three. Pick who, on purpose.
The modelling is standard but the data engineering is not โ harmonising six satellite products with different native grids, projections, temporal sampling and missing-data conventions onto a common daily quarter-degree grid is where most of the work sits, and teams consistently underestimate it.
Demo-ability
EasyEasy to demo โ and so is everyone else's. Working is the floor here, not the achievement.
A reconstructed vertical temperature profile laid over what a float actually measured is immediately convincing, and a depth-resolved section of the Bay of Bengal is visually striking in a way tabular results never are.
Data
None suppliedNo dataset comes with this one, so every accuracy figure you quote is a number about labels you invented.
Nothing is provided with the statement. You are sourcing, cleaning and labelling it yourself, and that work is invisible in the demo but very visible in the questions.
The demo they will have already seen
Somewhere around 95โ220 teams are heading here, and the description is doing the choosing for most of them. They will read the same brief, reach the same architecture, and build a version of the same demo you are planning. Being correct is the floor. If your five minutes could be swapped with the team before you and nobody in the room would notice, you have not picked badly โ you have built predictably, which costs exactly the same and hurts more.
What survives
The ground worth standing on when the questions start.
- Validation is against independent float observations rather than a held-out portion of the training product, which is a genuinely rigorous setup and lets you make a defensible skill claim instead of a self-referential one
- The domain, grid, inputs and output depths are all fixed by the statement, so you cannot lose time on scoping and cannot be criticised for a convenient choice of region or depth range
- The physical reasoning is sound and citable โ sea surface height responds to thermocline displacement, so there is genuine information about the subsurface in the surface fields and you can explain why the model works rather than only that it does
Nothing here is fatal. It is just the list of places this statement pushes back, and you now get to push there first.
The framing is a joke. The findings are not โ they are the same analysis on the statement page, and every line above is attached to a score or a fact in the record. It is one opinion with its reasoning attached, so argue with it before you trust it.