Scraping a Government Weather Site Because There's No API
Mauritius Meteorological Services publishes accurate sunrise, sunset, moonrise, and moonset tables on a public website — with no API. Here's the scraper, the cache, and the fallback.
metservice.intnet.mu/sun-moon-and-tides-moonrise-moonset-mauritius.php is not an API. It's an HTML table published by the official Mauritius Meteorological Services site, updated monthly, with no JSON endpoint, no rate-limit headers, and no terms of service page that mentions programmatic access at all. It is also the only source of accurate moonrise and moonset times for Mauritius — the closest paid API covers latitudes that don't match local conditions well enough to matter for solunar fishing theory. So the integration is a scraper, and treating a government page respectfully is a design constraint, not an afterthought.
Respectful scraping: cache aggressively, fetch monthly tables, never poll on every request.
Two data sources, two extraction methods
Sunrise and sunset come from the meteomoris Python library, which already wraps the sunrise/sunset table on the same domain — no need to write that parser ourselves. Moonrise and moonset aren't covered by that library, so that half is a custom BeautifulSoup scraper:
def get_moonrise_moonset(month, year):
url = f"https://metservice.intnet.mu/sun-moon-and-tides-moonrise-moonset-mauritius.php"
response = requests.get(url, timeout=10)
soup = BeautifulSoup(response.content, "html.parser")
table = soup.find("table", class_="tablepress")
return parse_monthly_table(table, month, year)The Node backend doesn't run this directly — it spawns the Python script as a child process and reads JSON off stdout:
const python = spawn('python3', [pythonScript, 'moonrisemu']);
let output = '';
python.stdout.on('data', (data) => { output += data.toString(); });
python.on('close', () => {
const result = JSON.parse(output);
resolve(result);
});Cross-language spawning is not the elegant choice — it would be simpler to have one runtime. It's the pragmatic one: BeautifulSoup and the existing meteomoris library are both Python, and rewriting a working scraper in cheerio for the sake of a single-language backend wasn't worth the risk of introducing new parsing bugs in code that's already working.
Why the cache duration is the whole ethics conversation
The site publishes data a month at a time — "sunrise for every day in February" comes back in one response. That means the correct caching strategy isn't "cache until the number changes," it's "cache until there's any reason to refetch at all," which for monthly tables is generously long. The integration caches for 24 hours in memory:
const CACHE_DURATION_MS = 24 * 60 * 60 * 1000;
let cache = { data: null, fetchedAt: 0 };
async function getCachedMoonTimes() {
if (cache.data && Date.now() - cache.fetchedAt < CACHE_DURATION_MS) {
return cache.data;
}
cache.data = await scrapeMoonTimes();
cache.fetchedAt = Date.now();
return cache.data;
}One fetch per day, regardless of how many users log trips or view the landing page's live conditions widget that day. Every user of Fishing Tracker Pro shares one scrape. That's not just polite to a government server that wasn't built for concurrent load — it's the only caching policy that makes sense given the data only changes once a month anyway.
The fallback is not optional
Government infrastructure goes down, gets redesigned, or blocks scrapers without warning, and a fishing app that goes dark every time a met office website hiccups is not a serious product. So there's a second, entirely independent code path: a mathematical approximation based on a seasonal sine wave tuned for roughly 20°S latitude, used automatically whenever the scrape fails or times out:
function approximateSunriseSunset(date, latitude) {
const dayOfYear = getDayOfYear(date);
const declination = 23.45 * Math.sin((360 / 365) * (dayOfYear - 81) * Math.PI / 180);
// ... seasonal offset applied to a base 06:00/18:00
}It's less accurate than the scraped table — by minutes, not hours — but it means the app degrades gracefully instead of failing outright. Users logging a trip during a met-office outage get slightly approximate solunar timing rather than a blank field or a 500 error.
Series: Fishing Tracker Pro. Next: how these environmental signals feed a trip-recommendation engine that suggests when to go fishing, not just where.