Seba — a Personal Guide at Museums

By CueGroup

An Adaptive Spatial & Contextual Intelligence Platform for Museum and Park Venues.


1. Executive Summary

We are building an AI based "personal guide in your pocket" platform. This is a conversational, multilingual AI agent that functions as a guide for people that are visiting venues like museums, national parks, zoos, etc.  It recognizes exactly what a visitor is looking at (an exhibit, a landmark, an animal, a machine on a factory floor) and delivers clear personalized commentary through the visitor's own mobile phone and headphones.  It then answers follow-up questions in natural conversation.  Additional explanatory information can be viewed on the phone screen.

Unlike legacy audio-guide apps (fixed-number keypad tours, GPS-triggered clips, static scripts), our platform is built around three differentiators:

·        Multi-modal, high-precision "what am I looking at" detection — computer vision as the primary sensor, backed by UWB, NFC, QR, and inertial gyroscope reckoning for the ~10–20% of cases where vision alone is ambiguous (like crowded galleries, glass reflections, look-alike specimens).

·        A conversational agent, not a script — visitors can ask anything, in their own language, and get an answer grounded in a deep knowledge base that extends far beyond what's physically on display.

·        Adaptive personalization — the system infers a visitor's interest level, background knowledge, and pace, and continuously adjusts vocabulary, depth, and topic selection, the way a competent human docent would.

The same underlying platform serves museums, national parks, zoos, botanical gardens, historic cities, and industrial/factory tours — any setting where a self-paced visitor benefits from context-aware, on-demand expertise. 

Seba means ‘to teach’ in ancient Egyptian.  The goal here is to teach individuals by offering accessible information. The star hieroglyph represents the wayfinding guiding light.  Also, the ancient Egyptian word for the hieroglyph sounds like ‘seba’.


2. The Problem

Cultural and tourism venues face a persistent trade-off:

·        Human docents are expensive, don't scale, aren't multilingual, and vary widely in quality — most venues can only offer scheduled group tours, not always-available 1:1 guidance.

·        Traditional audio guides (numbered keypad devices or GPS-triggered apps) are one-directional, one-size-fits-all, and expensive to produce and maintain. Content is scripted months in advance and can't answer a visitor's actual question.

·        Printed/static signage shows the same shallow information to the eight-year-old and the subject-matter expert alike, and visitors report attention spans of well under a minute per object.

·        Institutional collections are mostly invisible. A typical museum displays only ~10% of its holdings; the other 90%, plus related items at peer institutions, are curatorially rich but practically inaccessible to the general visitor.

·        Venues have almost no visibility into what actually interests their visitors, which exhibits under- or over-perform, or where visitors get confused or drop off.

The result: visitors get a shallower, less personal experience than they want, and venues lack the tools to fix it or to monetize deeper engagement (memberships, gift shop, repeat visits, donor cultivation).


3. The Solution

A cloud-based conversational AI "guide" delivered through a lightweight mobile app (or the venue's existing app or web page) and the visitor's own headphones and screen. At a high level:

Visitor experience:

·        Visitor opens the app (or web page), picks a language and (optionally) a quick interest/knowledge self-assessment. The app deduces the visitor’s age group, sex and general appearance using the device’s front camera to form an initial estimate of the visitor’s interests.

·        As they walk through the space, the phone's camera silently and continuously identifies the object, artwork, animal, machine, or landmark in view.

·        Headphone audio begins automatically — a short, high-quality narrative pitched to that visitor's inferred interest level — with no button-pressing required.  The phone’s screen may display additional information, like a video of how the object was used.  Audio information can be displayed as text for the hearing impaired.

·        The visitor can interrupt at any time: ask a spoken question, tap a suggested follow-up chip, or type a question, depending on venue norms (silence-sensitive spaces default to visual/text interaction).

·        The system remembers the conversation across the whole visit, so later answers build on earlier ones ("since you asked about the Buddha's hand gesture earlier, this next piece uses the same mudra...").

·        Groups (families, school classes, tour groups) can share one "session" so a single narration and Q&A stream serves multiple linked devices, with the group leader able to moderate or hand off the microphone.

Institution experience:

·        A content and configuration console for curators to ingest object records, wall text, catalog data, scholarly sources, and media — including for the 90% of the collection not on physical display, and (with partner agreements) related holdings at other institutions.

·        Analytics dashboard: dwell time, question topics, drop-off points, language mix, popular vs. neglected objects, conversion to membership/gift-shop prompts.

·        Configurable "voice," tone, content depth limits, and moderation rules per venue. 

·        And, of course, there could be ads that attract the visitor to the gift shop or cafeteria.


4. Use Cases

·        Art & history museums (e.g., a Southeast Asian Buddhist relics collection): deep object narratives, cross-referencing related pieces in storage or at partner museums, provenance and conservation stories, comparative religion/art-history context adapted to visitor background.

·        National / state parks: geo-triggered trail narration, wildlife and geology identification via camera, safety alerts, weather-aware suggestions, multilingual interpretation for international visitors.

·        City / heritage tourism: landmark recognition while walking, self-paced walking tours that adapt to how much time a visitor has left, restaurant/shop recommendations as a secondary revenue layer.

·        Zoos & aquariums: species identification, live-feeding-time alerts, conservation storytelling, kid-mode simplification vs. keeper-level detail for enthusiasts.

·        Botanical gardens: plant/species ID, seasonal bloom guidance, self-guided "theme" tours (medicinal plants, native species).

·        Factory / facility tours: safety-first narration, process explanation calibrated to visitor role (new employee vs. investor vs. student vs. supplier), equipment recognition on a moving line, group-tour audio sharing for plant tours.

The common thread: a self-paced visitor, a physical space with many points of interest, and a need for expertise that doesn't scale with human staff.


5. Benefits

For visitors:

·        a private, patient, infinitely knowledgeable guide that never rushes them, meets them at their level, speaks their language, and never runs out of answers.

For institutions:

·        Dramatically lower marginal cost of "deep" content versus scripted audio tours or additional docent staffing.

·        New monetizable engagement: premium tiers, membership upsell moments, gift-shop and event recommendations at natural conversational junctures.

·        Rich, first-party visitor analytics (what's engaging, what's confusing, where people spend time) that most venues currently cannot get at all.

·        A path to make the 90% of the collection that's in storage newly "visible" and valuable to visitors and researchers.

·        Accessibility gains: text/visual fallback for visitors who can't or shouldn't speak aloud, and language coverage most venues could never staff for.

·        Raise security alerts if a visitor wanders into restricted areas


6. Competitive Landscape

The category is active but fragmented, and most existing players are closer to "smart audio guide" than "conversational AI agent":

·        AITourPilot (Spain-headquartered) — AI-powered storytelling and personalized audio tours for museums, with a visitor-analytics and revenue (membership/gift-shop) angle; delivered via QR/link rather than continuous camera recognition. Reportedly early-stage and not yet institutionally funded.

·        Musa Guide (England) —  positions itself explicitly as conversational, multilingual, and built "from scratch" rather than as an automated version of a scripted tour.

·        Guru Experience Co. (San Diego, CA) — an established (~decade-old) museum app platform (e.g., San Diego Museum of Art, USS Midway) offering audio tours, wayfinding/beacon positioning, AR, and a CMS; recently adding AI options, but the core product line is a traditional branded-app-plus-beacons model rather than an always-on conversational agent.

·        Hearonymus (Austria) — a broad, multi-venue marketplace app for downloadable, offline audio guides across museums and cities; strong distribution and offline support, but content is largely traditionally produced/scripted per venue rather than a real-time, camera-driven conversational system, and app-store reviews flag UX and pricing-transparency issues.

·        Cuseum (Boston) —  a well entrenched company that provides digital and some AI services to many museums in the USA.  It does not provide a “guide” service for visitors, but is listed as a competitor because it could block access to some museums.

·        Adjacent players worth monitoring: various AI landmark-scanner apps (e.g., consumer "point your camera at a object" apps, like Google Lens, WalkNicely, iNaturalist, etc) that target individual travelers rather than institutional partnerships.

Our differentiation:

We combine

·        continuous, largely hands-free object recognition via a fused sensor stack rather than QR/keypad/GPS triggers alone,

·        a genuinely conversational, AI-RAG-grounded Q&A layer rather than a fixed script,

·        real personalization/adaptive difficulty rather than a single track per audience segment, and

·        a knowledge base deliberately extended beyond the display floor.

Execution risk is real — several competitors have years of institutional relationships and content-production pipelines already in place — so a credible go-to-market likely means partnering with, or acquiring content/relationship advantages from, an existing player rather than cold-starting museum relationships from zero.


7. Core Technical Capabilities

7.1 Object / location recognition (the hard problem)

A layered detection stack, from cheapest/most-available to most-precise, fused via a confidence-weighted model:

Layer

Role

Strength

Limitation

Computer vision (on-device + cloud)

Primary identification of the specific object/artwork/specimen in the camera frame

Works with zero venue infrastructure; handles fine-grained recognition (this statue vs. that statue)

Struggles with glare, occlusion, crowding, near-duplicate objects, low light

QR / NFC tags

Deterministic fallback and cold-start disambiguation

100% accurate when scanned; cheap to deploy

Breaks the "just look and listen" magic; requires visitor action; tags can be damaged/removed

UWB (ultra-wideband) beacons

Sub-meter indoor positioning to narrow the candidate set before vision confirms

Very precise, works without camera pointed correctly

Requires infrastructure investment (anchors) and device UWB support

Gyroscope / IMU dead-reckoning

Tracks visitor orientation and motion between fixes

Cheap, always-on, smooths out gaps between other signals

Drifts over time; needs periodic re-anchoring

GPS / Wi-Fi RTT

Coarse zone-level location (which gallery, which section of the park)

Free, works outdoors and roughly indoors

Meters-level accuracy indoors — not sufficient alone for object-level ID

 

The system is designed so a venue can launch with vision + GPS/Wi-Fi alone (near-zero infrastructure cost) and progressively add UWB anchors or QR tags in high-density or high-value areas to raise accuracy.

7.2 Conversational, multilingual AI

Speech-to-text and text-to-speech in a broad set of languages, with natural-sounding, low-latency voice output.

A retrieval-augmented generation (RAG) architecture: responses are grounded in a curated, venue-approved knowledge base (not open-ended web generation), which is critical for factual accuracy in an educational/cultural setting.

Fallback UI (suggested-question chips, text input) for quiet-zone etiquette or accessibility needs, so the experience degrades gracefully rather than requiring speech.

7.3 Personalization / adaptive difficulty

A lightweight visitor model built from an initial assessment based on the visitor’s image (determines the age group, sex, and ethnic background).  The visitor may also speak and set some initial constraints (like time or location limits, specific areas of interest, whether it is a group, etc.). The model is then continually refined by engagement signals (dwell time, follow-up question complexity, requests to "go deeper" or "keep it simple"), and vocabulary choices the visitor makes.

The narration engine selects content depth, vocabulary, analogies, and even which stories to tell (e.g., art-historical vs. narrative/legend vs. conservation-science framing for the same object) based on this model, and adapts continuously through the visit.

The visitor’s image can be saved so that the session can be resumed at a later visit.  All the personal attributes that were collected earlier can be restored.  With consent, this restoration can also happen at other venues that also use this system.

7.4 Knowledge base beyond the gallery floor

Ingests full collection records (not just the ~10% on display), museum archives, scholarly catalogs, and — via partnership agreements — related holdings at peer institutions, so a visitor asking "are there other pieces like this?" can get a real, sourced answer rather than "I don't know."

Content pipeline includes curator review/approval workflows so venues retain editorial control over what the AI says about their collection.

7.5 Architecture

Thin client (phone/tablet app) handles camera capture, audio I/O, and local sensor fusion; heavier inference (vision matching, LLM reasoning, TTS) runs on a cloud backend reached via the visitor's Wi-Fi or carrier data.

Session state (conversation history, visitor model, group session) lives server-side, keyed to a session token, not a persistent identity, by default.

SDK/white-label option lets venues embed the guide inside their own existing app rather than requiring a separate download.

More details of the implementation architecture are described separately, as is the project plan.


8. Business Model

·        White-Label SaaS: license the platform to venues (museums, parks, zoos, cities, facility owners) on a per-visitor or per-site subscription basis, similar to existing audio-guide vendors but with materially higher usage-based ceilings given richer engagement.

·        Content/onboarding services: fee-based ingestion and curation of a venue's collection/knowledge base, especially the "hidden 90%."

·        Revenue share: on membership conversions, gift-shop referrals, or ticketing upsells surfaced conversationally.

·        Consumer add-on tier: optional direct-to-traveler subscription for cross-venue access (city + museum + park in one trip), similar in spirit to Hearonymus's marketplace model but with the conversational layer as the premium differentiator.


9. Key Risks & Open Problems

·        Recognition accuracy at scale: fine-grained visual distinction between similar objects (e.g., near-identical Buddha statues, similar tree species, similar aircraft parts on a factory line) is a hard computer-vision problem; false identifications directly damage trust and require careful confidence thresholds and graceful fallback (e.g., prompting "which of these three are you looking at?").

·        Factual accuracy / hallucination risk: an AI "getting it wrong" about a religious or culturally sensitive object is a reputational risk for both us and the venue; requires strict RAG grounding, source citation, curator sign-off workflows, and conservative behavior when confidence is low.

·        Cultural and religious sensitivity: for a Buddhist relics collection specifically (and similarly for indigenous, sacred, or contested-provenance items elsewhere), content must be reviewed with subject-matter and community input, not generated freely — this is a curation/partnership requirement, not just an engineering one.

·        Connectivity dependence: the cloud-inference architecture assumes reliable Wi-Fi/carrier data; many museum basements, park backcountry, and factory floors have poor coverage, requiring either venue Wi-Fi investment, edge/on-device fallback models, or pre-cached content for known dead zones.

·        Infrastructure cost for high-precision positioning: UWB anchor networks are effective but require capital investment and installation per venue, which may limit adoption among smaller or budget-constrained institutions.

·        Privacy and data governance: continuous camera use and voice input raise legitimate visitor privacy concerns (especially involving minors); requires on-device processing where possible, clear consent flows, data minimization, and compliance with regional regulations (GDPR, COPPA, etc.) — likely a first-order product and legal workstream, not an afterthought.

·        Hardware and etiquette constraints: not all visitors have headphones, are comfortable speaking aloud in a quiet gallery, or want another screen between them and the art — the multi-modal fallback (chips/text) mitigates but doesn't eliminate this. The venue may need to rent out devices with the necessary hardware capabilities.

·        Competitive and incumbency risk: established vendors (Guru, Hearonymus, and others) already hold multi-year venue relationships and are adding AI features; the fastest path to scale may be partnership, integration, or acquisition rather than pure head-to-head competition for new venue logos.

·        Group-mode complexity: reliably sharing one audio/mic session across multiple linked devices (for families or tour groups) without confusing turn-taking or audio conflicts is a non-trivial UX and systems problem.

·        Content production cost/scale: even with AI assistance, ingesting and curator-approving deep content for 90%+ of a large collection, plus partner-museum holdings, is a significant content-operations undertaking and an ongoing cost center.

·        AI computing expense: The cost of using AI can be high.  Tokens are expensive the demand for AI is growing.  However, hyperscalers are expected to address the demand and drive down costs.

·        Monetization ceiling at smaller venues: many potential customers (small museums, local parks) have limited budgets, which may push the model toward larger institutions and city/park-authority contracts first, with a lighter self-serve tier for smaller sites.


10. Summary

This is a platform play, not a single-venue app: the same recognition, conversation, and personalization stack generalizes across museums, parks, zoos, gardens, cities, and industrial tours, with venue-specific content and infrastructure layered on top. The core technical bets — multi-modal recognition fusion, grounded conversational AI, and real-time personalization — are achievable with current technology, and the addressable market spans cultural, tourism, and industrial sectors globally. The primary execution risks are content operations, positioning-infrastructure cost, and competitive relationship-building with venues that already have incumbent vendors — all addressable with focused go-to-market sequencing (start with 1–2 flagship venues, likely a cultural institution with a rich "hidden collection" story, to prove the model before horizontal expansion).

Aug 11, 2026