PassportAI is a regulatory research service by itsaura for compliance departments of financial institutions in the EU, the EEA and the United Kingdom. It answers questions about EU law and the national rules of home and host states, with every statement linked to its public source. The crawler collects those sources.
It does not collect personal profiles, does not log in, does not submit forms and does not fetch anything behind a login or paywall. The automated crawler does not pass CAPTCHAs (see the supervised session below for the one exception). Content is used to answer research questions with citations to the original source.
Crawl-delay. We follow the group that names PassportAI-research, and otherwise
the User-agent: * group (RFC 9309). Groups that only name other crawlers (for example AI-training
crawlers) do not apply to us. If robots.txt cannot be fetched because of a server error, we fetch nothing from
that host in that run; a 404 means no restrictions. Text-and-data-mining reservations are respected as well.
Disclosed exceptions are listed below.Supervised browser session (separate from the crawler). A small number of public regulator sources that sit behind a CAPTCHA or browser check are fetched at most once a month in a visible desktop browser, operated and watched by a person at itsaura, from a different (office) connection. Where an access check appears, that person completes it by hand; nothing solves CAPTCHAs automatically and no check is bypassed. In that session the browser identifies as an ordinary desktop browser, not as PassportAI-research, and its requests are not signed. It follows the same rules otherwise: robots.txt per path, a slow human pace (at least six seconds between pages on a host), and only public documents. Sites that would rather not be visited this way can tell us through the Support button at the top of this page, and we will stop.
On the explicit decision of itsaura, and only for these paths, the crawler fetches material that the site's robots.txt closes for general crawlers. Each is official public material, fetched slowly, and each publisher has been or is being informed; we stop on request. The material is used for reference only: PassportAI cites and links to the original source and does not republish it as a dataset, and the site owners have been notified.
www.ejustice.just.fgov.be/eli/ — Belgian legislation (Justel, ELI pages).api.lovdata.no — only the two public NLOD bulk files of Norwegian legislation, at most once a day.sedlabanki.is/library — the Central Bank of Iceland's publication library (rules and guidance).www.cysec.gov.cy/en-GB/legislation/, /en-GB/public-info/, /Files/Law-Documents/
(also under /en-GB/ and /el-GR/), /CMSPages/GetFile.aspx and
/getattachment/ — the Cyprus Securities and Exchange Commission's directives, practical guides,
circulars, policy statements, announcements, Board decisions (fines) and the court judgments it publishes; monthly.www.cylaw.org/cgi-bin/open.pl and /apofaseised/index_ — CyLaw judgments on banking,
capital markets, insurance, anti-money laundering, sanctions, data protection and bribery (the year lists are read
only to select those cases); monthly.www.mof.gov.cy/mof/gpo/gazette.nsf/ — the Official Gazette of the Republic of Cyprus, Annex III
Part I (regulatory administrative acts), only the acts in the financial, anti-money-laundering, sanctions,
data-protection and anti-bribery areas; monthly.www.supremecourt.gov.cy/judicial/sc.nsf/ — judgments of the Supreme Court of Cyprus on banking,
capital markets, insurance, anti-money laundering, data protection and bribery; monthly.www.law.gov.cy/law/mokas/mokas.nsf/ — MOKAS (the Cyprus financial intelligence unit): reporting
guidance, typologies, strategic analyses and annual reports; monthly.odluke.sudovi.hr/Document/ — Croatian court decisions (higher courts, anonymised, since 2015) on
banking and consumer credit, insurance, capital markets, anti-money laundering, data protection and bribery; monthly.www.haod.hr/Portals/0/ — the Croatian Deposit Insurance Agency's rules, premium methodology and
annual reports; monthly.www.rechtsprechung-im-internet.de/rii-toc.xml, /jportal/docs/bsjrs/ and the document
pages under /jportal/ — federal court decisions (BGH, BAG, BVerfG) on banking,
capital-markets, insurance, payment, anti-money-laundering, data-protection and bribery law since 2016; the
site's robots.txt admits only the EU ECLI crawler; official works under § 5 UrhG; reference use only;
publisher notified.www.scj.ro (High Court of Cassation and Justice of Romania) — decisions of the administrative and tax
chamber involving the ASF or the BNR and decisions on the main Romanian financial-sector laws, at most once a month,
several seconds between requests; the pages carry a “noindex” instruction; reference use only; the court
has been notified.www.rv.hessenrecht.hessen.de (Bürgerservice Hessenrecht) — decisions of the VG Frankfurt am Main,
the Hessischer VGH and the OLG Frankfurt am Main since 2016 on banking, capital-markets, insurance, payment,
investment-fund and anti-money-laundering law or involving BaFin, fetched through the portal’s public guest
access (the same anonymous session every visitor receives) and its search, at least three seconds between
requests, at most once a month; the site’s robots.txt excludes general crawlers; official works under
§ 5 UrhG; reference use only; the Hessian Ministry of Justice has been notified and asked for a preferred
route.Requests carry Signature-Agent, Signature-Input and Signature headers
(Ed25519, tag web-bot-auth, covering @authority and signature-agent). A request
that claims to be PassportAI-research but comes from another address and carries no valid signature is not ours.
Add this to your robots.txt to block the crawler completely:
User-agent: PassportAI-research Disallow: /
Or ask for a slower pace:
User-agent: PassportAI-research Crawl-delay: 10
Changes are picked up on the next run. You can also contact us through the Support button at the top of this page and we will exclude your site or paths by hand.
Use the Support button at the top of this page (category "Other") and mention the host name and, if possible, a timestamp from your logs. We answer on working days and stop crawling a site immediately on request.