About the role
Indian trade regulation moves constantly. A duty remission rate is halved in February and restored in March. A notification supersedes another notification that amended a third. Nobody sends you an email about it, and an exporter finds out at the port.
You build the system that watches, diffs, extracts and routes — so a change in a circular becomes a machine-readable delta with an owner and a deadline, hours after it is published.
What this job entails. Ingestion across a long tail of government sources that were never designed to be parsed. Structural diffing that says "this clause changed, in this way, with effect from this date" rather than "this page changed." Extracting obligations from prose into something typed. Mapping an obligation to the specific customers it actually affects, so an alert is about them rather than about everyone. And temporal queries, because "what were the rules on the date of this shipment" is a different and much harder question than "what are the rules."
This is the asset that compounds. Products pivot; a versioned, diffed corpus of trade regulation survives all of it and is worth more every month.
One thing we will say plainly because it comes up: we are not looking for sentiment analysis. Regulation carries obligation, not opinion, and a sentiment score over a government circular is both noise and impossible to audit. The techniques that matter here are change detection, information extraction, entity resolution and temporal reasoning. If you want to argue with that, do — it would be a good conversation.
You do not need to know trade law. You do need to care that an extracted fact without a citation and a date is worse than no fact at all.
What you’ll do
- Build and run ingestion across government sources — notifications, circulars, tariff schedules, restricted-party lists.
- Build structural diffing that identifies what changed at clause level, and from when.
- Extract obligations from prose into typed records that carry their citation.
- Map obligations to the customers they affect, using their commodity codes, ports and destinations.
- Build temporal query support so the system can answer as of a past date.
- Own data quality, because everything downstream inherits it.
What we look for
- Roughly three to seven years building data pipelines in production.
- Strong Python and SQL, and comfort with the messy end: scrapers that break, PDFs that are images, ten different date formats.
- You have done information extraction, change detection or entity resolution on real data, not in a notebook.
- You treat provenance as a requirement rather than a nice-to-have.
- You can run something that has to be right every day, not just once.
Nice to have
- Regulatory, legal or government text specifically — horizon scanning, compliance content, legal publishing, or a regtech product team.
- Document AI and layout-aware extraction.
- Temporal or bitemporal data modelling; knowledge graphs.
- Multilingual extraction.
- You have worked somewhere the extraction had to be defended to an auditor.