Skip to content
SIH Buddyby Ganeev Singh
Dev
All problem statements
SIH26167Worth consideringacceptance 3/5

SatQuery AI - An Interactive Vision-Language Assistant for Multimodal Remote Sensing Image Analysis through Text Queries

Indian Space Research Organisation(ISRO) · Space Technology · Software

The datasets are handed to you and the optical-SAR angle is genuinely valuable, but VLM fine-tuning is compute-heavy and the multisensor reasoning is what tends to get dropped — commit to the paired and cross-sensor cases, since a single-image VQA model is the crowded easy version.

Data: BigEarthNet (Sentinel-1 SAR + Sentinel-2 optical); VRSBench, RSVQA for evaluation

What it actually is

Non-experts cannot easily get answers out of satellite imagery because it requires knowing GIS workflows, sensors and task-specific models. The ask is a vision-language assistant fine-tuned for remote sensing that answers natural-language questions about satellite images, including harder cases needing multiple images — optical plus SAR, or before-and-after pairs for change detection.

What to build

A remote-sensing vision-language assistant that takes a natural-language query and one or more satellite images and answers it, domain-adapted so it understands sensor characteristics and remote-sensing terminology rather than relying on a general VLM, using BigEarthNet's co-registered Sentinel-1 SAR and Sentinel-2 optical data for the multisensor adaptation, handling single-image questions, multitemporal change queries across image pairs, and fused optical-SAR reasoning where SAR adds structural and all-weather information, possibly orchestrating specialised task models behind the language interface, evaluated on VRSBench and RSVQA.

Smallest thing that wins the room

Ask the assistant a change-detection question over a before-and-after Sentinel pair — how has built-up area changed here — and get a grounded answer, then ask a question answerable only by fusing the optical and SAR views to show the multisensor adaptation working.

How crowded this one gets

A guess, projected from the 2025 statements — the last year where both the submission counts and the winners were published.

Moderate90–210 teams expectedroughly 1 in 76–176 wins it

Quieter than 70% of the 226 · #69 of 226 by expected field

A normal-sized field. Your idea has to be good, not miraculous.

Why: defence, intelligence and space bodies drew small fields.

This is a guess, not a fact

Nobody has published 2026’s numbers yet. This is an analysed estimate from last year’s pattern, so please do not take it as the truth — check the live counter on the SIH portal before you decide anything. The range covers the middle half of likely outcomes, so one statement in two lands outside it. Entry closes at 500 ideas per statement, so no range goes past that — a statement that reaches the cap fills and shuts rather than drawing an unlimited crowd. The model reads only three things a team can see before choosing — software or hardware, the theme, and what kind of body posted it — and those explain about a quarter of the variation in last year’s field sizes (R² 0.25 on held-out statements). Trust the band more than the number, and the ordering more than either. It cannot see how good your idea is, which is the part that actually decides it.

The scores

The number is the shorthand. The line under it is the reason.

What you will be writing

  • Remote-sensing VLM fine-tuning (LLaVA / Qwen-VL base)
  • BigEarthNet Sentinel-1 + Sentinel-2 co-registered data
  • Optical-SAR multimodal fusion
  • Multitemporal change reasoning
  • VRSBench / RSVQA evaluation
  • LoRA / parameter-efficient fine-tuning
  • Remote sensing
  • Vision-language models
  • Multimodal AI

Prior art to read before you start

remote-sensing VQA · optical-SAR multimodal reasoning · multitemporal change interpretation via language

Analysed by Claude Opus. Every score above is a judgment call with its reasoning attached — kindly cross-check this against the official statement on the SIH portal before your team commits to it.