Integrated Polar Science Outreach, Knowledge Repository and Media Dissemination Portal
Ministry of Earth Sciences (MoES) · Space Technology · Software
The corpus is real and the build is safe, but an archive with an AI post generator has a very low ceiling — take it only if you will make grounded, citation-locked generation and multimodal search the substance rather than shipping a CMS.
What it actually is
Decades of Indian Antarctic expedition reports, datasets, photographs and papers sit scattered across drives and old websites, and almost none of it reaches the public. The ask is one archive holding all of it properly, and a way to turn that material into things people actually read — website posts and social media content — without someone writing each one by hand.
What to build
A repository plus publishing pipeline: an ingestion and cataloguing layer for the material types the statement lists — expedition reports, scientific datasets, publications, photographs, videos and institutional activity records — with metadata extraction, expedition-and-year organisation and full-text plus semantic search across the corpus; and a content generation layer that takes a selected item or theme and drafts website copy and social posts at a stated reading level with the source material cited, an editorial review queue so nothing publishes unreviewed, and a scheduling view for campaign planning.
Smallest thing that wins the room
Search the archive for a specific expedition, open a report from it, and generate a short public-facing post that draws only on that document with the passages it used highlighted in the source.
How crowded this one gets
A guess, projected from the 2025 statements — the last year where both the submission counts and the winners were published.
Quieter than 67% of the 226 · #76 of 226 by expected field
A normal-sized field. Your idea has to be good, not miraculous.
Why: central ministry statements sat below the average.
This is a guess, not a fact
Nobody has published 2026’s numbers yet. This is an analysed estimate from last year’s pattern, so please do not take it as the truth — check the live counter on the SIH portal before you decide anything. The range covers the middle half of likely outcomes, so one statement in two lands outside it. Entry closes at 500 ideas per statement, so no range goes past that — a statement that reaches the cap fills and shuts rather than drawing an unlimited crowd. The model reads only three things a team can see before choosing — software or hardware, the theme, and what kind of body posted it — and those explain about a quarter of the variation in last year’s field sizes (R² 0.25 on held-out statements). Trust the band more than the number, and the ordering more than either. It cannot see how good your idea is, which is the part that actually decides it.
The scores
The number is the shorthand. The line under it is the reason.
Acceptance potential
2/5A content repository with an AI caption generator is close to the lowest-ceiling software shape available — the archive half is a CMS and the generation half is an LLM call, so there is very little for a judge to reward beyond execution and the outreach framing does not create technical substance.
Feasibility
4/5Indian Antarctic expedition reports, station documentation and imagery are published openly enough to assemble a genuine corpus, and repository search plus grounded text generation are mature and well-supported — nothing here depends on data you cannot obtain.
Innovation scope
4/5At 213 characters the statement names what the portal holds and that it should generate content, and prescribes nothing about how — the retrieval design, the generation grounding and the editorial workflow are all yours.
Clarity
2/5Short, but unlike the neighbouring polar statements it does at least enumerate the content types and identify two distinct functions in archiving and generating, so you are not starting from nothing — though there is no audience, no scale and no measure of what good output would be.
Effort
HeavyIngestion across six media types, metadata extraction, a search layer, a generation pipeline, an editorial review queue and a scheduling interface is six components, and handling video and imagery properly is more work than the text side.
Demo-ability
MediumSearch over a real archive works reliably and grounded generation with highlighted sources is a decent moment, but a content portal generating social posts is a familiar sight in 2026 and nothing here will surprise a panel.
In its favour
- Green flag: Indian polar expedition material is genuinely published and downloadable, so unlike the neighbouring NCPOR statements your corpus is real rather than invented
- Green flag: Grounding generated posts to specific passages in a source report is a real and demonstrable discipline, and a portal that refuses to write anything it cannot cite is a defensible differentiator over freeform generation
- Green flag: Semantic search across photographs and video, not just text, is the harder and more interesting half and almost no competing team will attempt it
- Green flag: The Space Technology theme label hides an archive and outreach portal from anyone browsing by subject
Against it
- Red flag: An archive plus a text generator is among the least technically distinctive products you can build, and a panel will have seen several retrieval-and-generate submissions before yours
- Red flag: Generated science communication that subtly misstates a finding is worse than no post at all, so the editorial review queue is a requirement rather than a nicety and should be visible in your demo
- Red flag: Six media types is deceptively broad — video ingestion, transcription and indexing alone can consume the whole event if you let it
- Red flag: There is no measure of good outreach in the statement, so any claim that your generated content is effective rests on your own judgement and a judge may simply disagree
What you will be writing
- hybrid keyword and dense retrieval over the corpus
- Dublin Core / DataCite metadata cataloguing
- grounded generation with source span citation
- CLIP-based image and video semantic indexing
- editorial review queue with approval states
- Next.js public portal with scheduling view
- Digital archives and knowledge management
- Science communication
- Content generation
Prior art to read before you start
institutional knowledge repository with search · grounded content generation from source documents · media asset cataloguing and dissemination
Analysed by Claude Opus. Every score above is a judgment call with its reasoning attached — kindly cross-check this against the official statement on the SIH portal before your team commits to it.