For shop owners and webmasters
GadgifyBot
GadgifyBot is the program Gadgify uses to read the public product pages of Thai shops and record prices. This page describes what the program does today.
What GadgifyBot is and why it reads
Gadgify helps Thai buyers choose headphones and see whether today's price is good or worth waiting on. That needs prices from many shops and a price history. GadgifyBot reads those prices.
What it reads
- A shop's public product pages, over plain HTTP requests, for the price, the list price and the stock status in schema.org JSON-LD. If the JSON-LD has no price, it reads the page's price meta tags.
- If a price cannot be read, that check is recorded as having no data. It never guesses a price, and a price in a currency other than baht is not recorded as a baht price.
- A shop's sitemap and listing pages, to find product page URLs. It reads at most 500 links and at most 5 listing pages per source in one pass.
- A newly found product page is opened in a headless Chromium browser to keep the page title, the visible text (up to 20,000 characters), the product's JSON-LD and a full-page screenshot, so a Gadgify admin can review it before the product enters the catalogue.
What it keeps, and what it does not do
- The screenshot is kept in private file storage and deleted after 7 days. The page text is kept in our database. Both are for admin review only and are not published on the site.
- It does not save product image files on their own. The image address in JSON-LD is noted as a link only, but the screenshot shows the whole page, so it includes the pictures that page displays.
- It does not open Shopee or Lazada web pages. Prices from those two platforms come from their affiliate APIs once access is approved.
- The browser stays on the shop's own host. A redirect to another host is not followed, and images, video and fonts from other hosts are not loaded.
- If it meets a bot-check page (such as a Cloudflare challenge), it stops at that page. It does not try to pass it, use proxies, or pose as an ordinary browser.
User agent
Every request, whether read over HTTP or opened in the browser, sends this string. The version number changes as the program is updated. The link in brackets points to this page.
GadgifyBot/0.1.0 (+https://heygadgify.com/bot)robots.txt
GadgifyBot reads each host's robots.txt before requesting any page and follows Allow and Disallow rules, using the longest matching rule and supporting * and $. The token is GadgifyBot, matched without regard to case. If a GadgifyBot group exists it is used; if not, the * group is used.
User-agent: GadgifyBot
Disallow: /User-agent: GadgifyBot
Disallow: /cart/
Disallow: /*?session=User-agent: GadgifyBot
Crawl-delay: 10- robots.txt is remembered for 24 hours, so a change may take that long to apply.
- If robots.txt is missing (a 4xx answer), no limits apply. If it answers 5xx or cannot be reached, reading is treated as forbidden and checked again after 15 minutes.
- At most 512 KB of robots.txt is read; the rest is ignored.
Crawl rate and backoff
The values below are the ones set in the program today.
- Requests to one host are made one at a time, never in parallel.
- There is at least 1 second between two requests to one host. If robots.txt gives a longer Crawl-delay, that value is used.
- If the Crawl-delay is longer than 60 seconds, the host is not read, rather than waited on.
- Each product page Gadgify tracks is read once a day, every 6 hours in the 7 days before a major sale day, and every hour on the sale day. If a read fails, it tries once more about a minute later before recording it as missing.
- A request with no answer within 20 seconds is cancelled. A response body larger than 5 MB is not used.
- On a 429 or 5xx status, a timeout or a network error, it retries up to 3 tries in total (the first included). The wait doubles each time, starting near 1 second, up to 30 seconds, with a little random spread.
- If the server sends Retry-After, it waits that long. If that is longer than 60 seconds, it stops instead of waiting.
- Other 4xx statuses are not retried. Redirects are followed at most 5 times, and each one is checked.
AI opt-outs it honours
If a host's robots.txt says it does not want AI to read it, GadgifyBot reads nothing on that host. GadgifyBot does not use data to train models, but we treat this as a refusal anyway. Two signals are checked today.
- A User-agent group for a known AI bot that contains Disallow: /. The names checked areanthropic-ai, claudebot, claude-web, claude-user, gptbot, chatgpt-user, oai-searchbot, ccbot, google-extended, applebot-extended, perplexitybot, bytespider, amazonbot, ai2bot, cohere-ai, diffbot, imagesiftbot, omgili
- A Content-Signal: ai-train=no line.
Other signals, such as meta tags or X-Robots-Tag, are not checked yet.
IP addresses
GadgifyBot crawls from no fixed IP addresses, so we have no IP list to check against. A user agent can be faked. If you doubt that a request came from us, email us with the time and the URL requested, and we will say whether it was ours.
Shops that do not want to be read, or want to send data themselves
A shop that does not want to be read adds a Disallow to its robots.txt, as in the examples above. GadgifyBot stops once the 24-hour memory has passed. When a shop blocks us, we do not look for another way to read that page. A shop that wants its prices on Gadgify without page reading can send a product data feed or ask to work with us.
Contact
Questions, requests to stop, or corrections: write to this address.