- GLAMorous AI
- Posts
- GLAMorous AI TL;DR — June 2026
GLAMorous AI TL;DR — June 2026
The Scraping Problem: Who Owns Your Collections?
"Collections as data are not collections for scraping."
Welcome to this month's edition, The Scraping Problem, a roundup of how the GLAM sector is grappling with an uncomfortable reality: your digitised collections, built with public funding and held in public trust, are quietly becoming training data for commercial AI systems. No consent. No compensation. No consultation.
This month's theme is sovereignty: who gets to decide what happens to your collections, and how do you keep control of what you've worked decades to preserve?
🌟 Featured Reads
GLAM-E Lab – Are AI Bots Knocking Cultural Heritage Offline?
The report that changed the conversation. Late 2024 and early 2025 saw a wave of reports from GLAM institutions describing their servers being overwhelmed by AI training bots. The GLAM-E Lab survey of 43 institutions across Europe, North America, and Oceania found consistent patterns: increased traffic spikes, scraped collections, systems straining under load. Of 39 respondents experiencing traffic increases, 27 attributed it to AI bots. On December 4 alone, one institution recorded 11,329 searches from thousands of different internet addresses.
This is not a technical problem. It is a governance failure.
🏛️ The Infrastructure of Care
Camberidge University – Collaborative Digital Stewardship: Cambridge's ArCH Community of Practice
Cambridge's new AI for Cultural Heritage Hub (ArCH) offers a different model: a secure workspace where curators, researchers, and archivists work with AI tools on real collection challenges, not as subjects of external AI experimentation. The emphasis is on integrating expert knowledge into training datasets, on institutional control of the workflow, and on AI as a tool for your questions, not someone else's training data.
It is a modest but important distinction.
Colavizza, Jaillant & Mukhopadhyay – AI Preparedness Guidelines for Archivists
A grounded reflection on what "responsible AI" means in archival work: transparency about data use, institutional consent, and community benefit-sharing. The authors argue that responsibility is not a feature you add to a model. It is built into the workflows, governance structures, and institutional policies that decide whether your collections remain yours.
📜 When AI Works Well
Avagliano, Canuto & Martocci – Explainable AI, LLM, and Digitized Archival Cultural Heritage: A Case Study of the Grand Ducal Archive of the Medici
A careful study from Florence showing where generative AI genuinely helps: post-processing structured catalogue records to create visual finding aids that make early modern correspondence searchable and navigable. The critical insight: the AI works because the data is clean, the task is bounded, and archivists verify every output.
It is useful precisely because it is not magic.
⚖️ Rights, Remedies, and Machine-Readable Consent
Pandit, Lindquist & Krog – Implementing ISO/IEC TS 27560:2023: Consent Records and Receipts for GDPR and DGA
If you cannot prove you did not consent to your data being used for AI training, you have no legal ground to stop scraping. This technical standard provides the infrastructure for documenting and verifying consent at scale. Machine-readable consent records, audit trails, and transparent processing agreements are not bureaucratic overhead; they are the load-bearing walls of data sovereignty. Without them, your "open access" becomes implicit permission.
Schmid & Popple – Copyright in 2026: Clarification, Review and Reform
Litigation is accelerating across Europe. Hungary, France, the Netherlands, the Czech Republic, Germany, and the UK are all grappling with whether AI companies can train on copyrighted works without authorisation. The critical gap: decisions in one jurisdiction do not bind another. A model trained legally in the US under fair use can be deployed in the EU, where the rules are stricter. Scrapers exploit this regulatory fragmentation. Your defence depends on where you are.
💼 Licensing Models Beyond Binary Open/Closed
Dişli and Candela – Copyright and Licencing for Cultural Heritage Collections As Data
Most GLAM institutions publish collections as "open" or "restricted." This binary misses the middle ground. The real work is in machine-readable, interoperable licensing that specifies what each reuser (human researcher, AI trainer, commercial developer) can and cannot do. A Creative Commons license on a webpage is not machine-actionable. Collections as data require metadata-embedded licenses that API clients must read before accessing.
Benhamou and Dulong de Rosnay – Open Licensing and Data Trust for Personal and Non-Personal Data: A Blueprint to Support the Commons and Privacy
An alternative framework: data trusts. Rather than institutions managing access alone, a trust structure gives communities, researchers, and institutions shared governance over how collections are used. It does not require closing collections. It requires explicit consent pathways and benefit-sharing arrangements. This model is already being tested in cultural heritage and Indigenous knowledge contexts.
💰 The Real Cost of Defending Collections
BCG Report (via Vocal Media) – When the Archive Becomes a Liability: Digital Transformation for Museums
Here is what the scraping crisis exposes: institutions failing to modernise legacy infrastructure spend 70% of their IT budgets on maintenance rather than innovation. That leaves almost nothing for defending against bots, upgrading servers, or implementing access controls. The irony is sharp: your collections are valuable enough for commercial AI companies to scrape, but your institution cannot afford to protect them. Scraping thrives on underfunded infrastructure.
🌍 The Regulatory Fragmentation Problem
CEIMIA – A Comparative Framework For AI Regulatory Policy: Copyright and AI
The EU, UK, and US have adopted fundamentally different approaches. The EU AI Act (August 2025) requires transparency and restricts the deployment of models trained on copyrighted works without authorisation in EU territory. The UK's approach is weaker: no territorial restriction. The US relies on fair use, which is broader. The result: a scraper can train a model in California or Canada, then export it globally. One jurisdiction's protection becomes another's loophole.
Bruegel – The European Union is Still Caught in an AI Copyright Bind
Even within the EU, the fix is incomplete. The AI Act's transparency rules require companies to publish a "summary" of training data sources. But they can omit small publishers and non-English content if transaction costs are high. The result: training datasets become biased toward English-language, large-scale collections. Smaller GLAM institutions, particularly those in non-English regions, get scraped and erased from the transparency record.
Kretschmer et al. – Copyright and AI in the UK: Opting-In or Opting-Out?
The UK consultation on copyright and AI attracted over 11,500 responses. The government's response: no new law. Instead, a delay and a promise to assess options by March 2026. Meanwhile, AI companies are already operating in the regulatory grey zone. The UK's lack of territorial restriction means models trained abroad without UK copyright compliance can still be deployed to UK users. This is not neutrality. It is a choice to leave collections undefended.
❓ Big Question
If scraping bots can train models in jurisdictions with weak copyright law, export them globally, and exploit regulatory gaps - then whose job is it to close those gaps?
Is it individual GLAM institutions, fighting in local courts? National governments, acting in isolation? International bodies, moving at glacial speed? Or does sovereignty over collections require something we have not yet built?
💬 About
I'm Alfie, a researcher and archaeologist exploring where heritage, ethics, and AI meet. This digest keeps things short, critical, and useful—no jargon, no hype.
👉 Read or subscribe at glamorousai.beehiiv.com
👉 Send papers, reports, or ideas for July. The conversation is only starting.