Skip to content

Design language: Perapian, bukan wajah

Status: DRAFT v0.1 (2026-07-05, design session with Yose). Governs device/00 spec and the e-ink UI. House style applies.

Thesis

Personal AI gets a face. Communal AI gets architecture. The design ancestor of Aksara is not the smart speaker or the robot; it is the papan pengumuman, the jam dinding, and the perapian di tengah honai: the hearth. A hearth is alive (light moves), warm, non-anthropomorphic, gathers people, and watches no one. Every design decision derives from this: the device is a place in the room, not a person in the room.

The OLED-eyes idea was correctly killed (uncanny). Anything that resembles gaze is banned permanently: no eyes, no face, no camera on the node, no darting motion.

Presence rules (the anti-creepy contract)

  1. Motion at the pace of breath, never the pace of attention. Idle animation cycles are minutes, not seconds. Nothing darts, nothing tracks.
  2. Light is indirect. The halo is reflected off the wall behind the device (firelight logic). No visible emitters, no camera-like dots.
  3. Attention is shown by leaning, not looking: the mic array gives direction of arrival; the halo wave leans toward the speaker.
  4. It shows work, not thought: "Menyiapkan 3 surat untuk besok. Antrean kosong. Sinkron terakhir 14:02." The anticipation layer made visible is the honest version of "thinking".
  5. Presence sensing without identity: an mmWave presence sensor (LD2410 class, ~USD 5) lets Aksara know someone entered the room without any image existing. Greeting on arrival (halo bloom, optional spoken greeting during office hours) feels like an attentive colleague; it is technically incapable of knowing WHO. This is the spirit sense: the room notices you, nothing watches you.

Sound hierarchy (decided with Yose, 2026-07-05)

Voice is crucial and confirmed: Aksara converses in speech (full TTS subsystem per v4.1 §25). The hierarchy governs when which layer speaks:

  1. Silence + light for wake acknowledgment: someone says the wake word, the halo leans toward them. No sound. (The colleague who looks up when you call their name.)
  2. Tifa tap for the bell layer: agenda reminders, meeting time, queue calls, a document ready for signature, sync completed after a long offline stretch. The tifa is the office bell, not the voice. Recorded with permission from a Papuan player, credited (LEDGER).
  3. Voice for conversation: answering, reading a draft aloud, confirming an action. Speaks only in an active exchange or when a bell escalates to explanation.

Camera position v2 (PROPOSED 2026-07-05, deliberation with Yose ongoing)

Revised after Yose's counterargument (offices already run CCTV; event capture for bookkeeping, local only, never streamed, is a real need). The synthesis: reframe the camera from a SENSE (always available, watching) to a TOOL (invoked per task, leaves receipts).

  1. No always-on camera, no face recognition, ever. Recognizing citizens by face is sensitive-biometric processing under UU PDP, and it kills the trust sentence in every audiensi and dewan meeting. This part stays hard.
  2. Kamera saksi (capture-as-transaction), allowed: a shuttered camera module whose every capture is an aksara-cli tool invocation. Physical shutter closed by default. Capture requires a transaction context (bukti penyerahan bantuan, lampiran notulen rapat, dokumentasi serah terima), emits a signed audit record, and announces itself: tifa tap + halo flash at the moment of capture (the shutter-sound law). Photos stored locally, encrypted, retention per policy, never federated. Indonesian administration already demands exactly these evidence photos; Aksara does them with receipts.
  3. The trust sentence evolves and survives: "AI pemerintah yang tidak bisa mengintai: setiap foto meninggalkan bukti dan bersuara." A camera that cannot take a photo silently or secretly is categorically different from CCTV.
  4. Person-aware interaction (the actual want behind the camera ask) is delivered by the identity substrate below, not by vision.

Identity substrate: how Aksara knows people (without watching anyone)

WhoMechanismNotes
Enrolled staff (the 3-10 people who matter daily)On-device speaker recognition (voice embeddings, ECAPA-class, local, tiny)Voice is already the primary modality; enrollment is consented and revocable like fingerprints (v4.1 §23); enables "Selamat pagi, Ibu Maria" + per-person preferences
Staff, for approvalsFingerprint stempel buttonAlready spec'd
Citizens, remoteWhatsApp number binding + NIK verification flowThe number IS the session identity
Citizens, at the counterNIK verification per transaction; no biometrics retained
Anyone entering the roommmWave presence (no identity)Greeting bloom; anonymous by construction

The employee feel comes from voice identity + memory + rhythm: knowing that Monday is apel pagi, that Ibu Maria handles kesra letters, that the LPJ deadline is Thursday. Colleagues are known by what they remember, not by staring. What differentiates Aksara from a smart board or kiosk is exactly this memory + the background work, which is why the memory stack (research/2026-07-05_memory-stack.md) is load-bearing for the product story, not an internals detail.

Statline: the real-time agent status bar

A dedicated one-line region (bottom of panel, A2 partial refresh on e-ink, free on any faster display) showing the agent's live state, updated 1-2x per second while working, sparse when idle:

  • Working: "Menyusun draf surat domisili (2/3) . antre: 1 . tenggat LPJ: 3 hari"
  • Anticipating: "Menyiapkan 12 surat untuk musim raya (perkiraan naik 40 persen)"
  • Idle: clock tick + stempel suasana breath
  • Community/dev mode: kaomoji allowed in the statline (homage acknowledged)

The statline is the "di balik layar" region made continuous: proof of background work is a product feature (PNS Digital works while the office sleeps), and it is the always-true answer to "what is it doing right now".

Form: Papan + Punggung (board + spine)

Asymmetry with a function. Landscape rectangle in two zones:

  • Papan (the board): the 10.3 inch e-ink panel, offset to one side, thin wing (target 15-18 mm visible edge). The reading zone. Layout idiom is a printed page or notice board, never a dashboard.
  • Punggung (the spine): a vertical column on one edge, deeper (target 30-34 mm), carrying everything with volume: Orange Pi 5 on the aluminum backplate, battery, speaker chamber, mic array slots, fingerprint stempel button, LED channel. The tactile zone: every physical interaction lives on the spine.

The asymmetry serves thinness: the visible panel wing stays thin because the spine carries the depth, like a clipboard with a pen rail. On the artisan edition the spine is the carving surface (motif under community consent, v4.1 roadmap).

Indicative envelope (to be verified in blueprint pass): approx 340 x 210 mm footprint (near A4 landscape), panel active area 209.7 x 157.2 mm (1872x1404, 227 dpi), wall mount on French cleat with 10-12 mm standoff (the halo needs wall distance; the standoff also hides the spine bulge), desk variant with fold-out easel leg.

The alive layer (e-ink content)

Monochrome 16-level grayscale is the medium; its constraints are the style.

  1. Lukisan hari ini: idle state shows a daily grayscale artwork. Sources, in order of preference: Papuan motifs and works licensed under CARE with a printed credit line ("Motif Kamoro, seizin dewan X"), school children's drawings scanned monthly, community photographs. The idle screen is a gallery the kampung curates; the device demonstrates the consent model daily.
  2. Slow handwriting: Aksara means script. The signature move is stroke-by-stroke handwriting animation (partial refresh, one stroke per second, e-ink's native rhythm): the date, a proverb in bahasa daerah with translation, the day's numbers. A custom hand ("tulisan Aksara") becomes the brand.
  3. Stempel suasana: status glyph as a small round stamp mark in one corner (office stempel metaphor), texture varies by state. Kaomoji reserved for community mode and dev builds, not government offices.
  4. Self-published SLA: documents this month, average processing time, queue state (specs/etnos/06).
  5. E-ink discipline: dithered images, no gray-on-gray text, full refresh on content change, partial refresh only for strokes/ticker/queue.

Halo spec

COB addressable strip in a rear channel with diffuser, warm amber-white palette only (firelight; never RGB rainbow). One-dimensional damped-wave animation (v4.1 §24). States: mendengar (slow rise, leaning toward voice), berpikir (gentle traveling wave), berbicara (amplitude follows speech envelope), butuh persetujuan (two soft pulses pooling near the stempel button), antre offline (slow amber breathing), darurat/mesh (distinct but calm pattern, device/04).

Placement and mounting (decided 2026-07-05): primary strip along the TOP edge, throwing light up the wall (the hanging-paper-with-light-behind feel Yose wants), with a short spill segment down the spine side used for direction-leaning. Geometry: French cleat + safety screw, 10-12 mm wall standoff, panel face tilted 3-5 degrees downward (standing-height readability, less glare).

Drunk-proofing (the last-day-of-office collision case): nothing protrudes. The strip and diffuser sit recessed INSIDE the rear standoff gap behind a frame lip; an impact hits the chamfered PETG frame and the wall, never the diffuser. The bezel stands 1 mm proud of the panel glass so face-on impacts land on frame, not glass. The cleat takes shear; the safety screw stops uplift. Field-replaceable strip on a connector, because someone will eventually manage it anyway.

Division of labor: the LED is the heartbeat, the e-ink is the mind

E-ink physically cannot loop continuous animation (waveform wear, ghosting, refresh flash). So continuous aliveness lives in the halo (LEDs exist for exactly this), and the e-ink pulses sparsely: a handwriting stroke, a clock tick, a stempel breath every 30-60 s. Stillness with a heartbeat reads as calm intelligence; a busy screen reads as a broken appliance. The panel also carries two standing regions: "di balik layar" (a small redacted live feed of what the agent is doing: "menyusun draf surat domisili"), and a self-explanation footer ("Apa ini? Aksara adalah asisten kantor ini. Data Anda tidak meninggalkan perangkat tanpa persetujuan."), rotating in id and bahasa daerah.

Color e-ink verdict (2026-07-05)

At 10.3 inch, color e-ink under USD 100 does not exist. Sub-USD-100 color exists only at ~7.3 inch (Spectra class) with 12-20 s refreshes and no fast partial mode, which kills the alive layer. Decision: mono 10.3 for grayscale art + fast partial refresh; the amber halo carries the color. Revisit color panels at Fase 2 prices. Panel line update per D19 (2026-07-11): the demo unit's mono 10.3 is the Seeed XIAO ePaper EE03 kit (ED103TC2 class panel + ESP32-S3 display node, USD 114.90 verified, ~Rp 2.6jt landed); the earlier IT8951 line (USD 203, ~Rp 3.9-4.4jt landed) is retired for the demo (research/2026-07-11_cheap-display-paths.md). The design verdict above is unchanged.

Internals and materials

  • Single-layer component layout on a 2 mm aluminum backplate (thermal path, v4.1 §24); no stacking; OPi 5 with low-profile passive plate coupled to the backplate.
  • Battery: 12Ah/153Wh LiFePO4 pack class (scout: ~Rp 1.1jt, better Wh per rupiah) if the spine depth allows; else 40-60 Wh flat pack per v4.1.
  • Prototype 1: FDM-printed frame (PETG, matte white/grey/black) + off-the-shelf aluminum sheet; no carpenter hunting this week. MJF nylon shell at pilot batch (JLC3DP quote is a LEDGER item). Wood veneer and carved spine are production finishes.

E-ink UI software stack (expanded 2026-07-05)

Verdict: Svelte + Python is the right stack; nothing more exotic buys features we need.

  • UI layer (JS/TS, Svelte): one static app, 1872x1404, grayscale design tokens, "e-ink simulation mode" toggle (grayscale filter, ghosting overlay, refresh flash) so any phone browser is a faithful dev target. The same app is the public demo and the device UI.
  • Render bridge (on device): headless Chromium (Playwright/puppeteer-core) screenshots the page on state change; Pillow quantizes to 16-gray with dithering (Floyd-Steinberg or blue-noise; hitherdither class libs) and diffs regions.
  • Panel driver: per D19 the ESP32-S3 display node in the Seeed EE03 kit drives the ED103TC2 panel and holds the page at microwatts; the OPi renders frames and hands them over (TRMNL-style seam; the IT8951 HAT path is retired for the demo, research/2026-07-11_cheap-display-paths.md). Waveform discipline unchanged: GC16 full refresh on content change (the flash is honest page-turning), A2/DU fast partial (~120-260 ms) for strokes, ticker, queue numbers; scheduled GC16 clear every N partials to purge ghosting. Partial-refresh quality through Seeed's stack is the TBV.
  • Animation assets: e-ink animation = choreographed frame sequences, not CSS transitions. Handwriting uses single-stroke (Hershey-class) fonts or custom recorded stroke paths: SVG stroke order rasterized offline into 1-bit frame stacks, played as A2 partial updates one stroke at a time. Daily art pipeline: any image, auto-dithered, credit line composited. Assets are generated by scripts, not hand-drawn per day.
  • Halo controller: the damped-wave engine is a tiny Python process feeding the addressable strip (SPI/PWM), listening to the same event bus as the UI; states are code, not assets.
  • Explicitly rejected: LVGL/Qt (rebuilds UI twice, no web reuse), Flutter custom embedder (complexity without benefit at 16 GB RAM), pure-Pillow UI (rebuilds layout logic the browser gives free).

Touch decision (2026-07-05)

No touchscreen on the node panel. Touch turns the hearth into a kiosk: it invites walk-up poking, demands menu UI, collects fingerprints on the daily artwork, and erases the differentiation from "smart board di kantor". Interactions stay voice + WhatsApp + stempel button. Exception: the MPP counter variant (portrait queue mode) is allowed a capacitive touch layer in Fase 2, because at a service counter kiosk-ness is wanted. One product, two stances: hearth on the wall, kiosk at the counter.

Access matrix: which surface does what

SurfaceWhoCan doCannot do
Panel (wall)everyone in the roomread: art, pengumuman, agenda, queue, statline, SLA, self-explanation; consent via stempel buttonno browsing, no touch, no transactions
Voice (in-room)enrolled staff (identified), guests (anonymous)converse, dictate naskah, ask status, trigger transactions, receive readoutsno approvals by voice alone (voice is convenience identity, not authority)
WhatsAppcitizens, officialsfull citizen transactions; official remote approval (liveness + passphrase); status checksnothing touching another person's data
Local web dashboard (LAN, office PC/phone)operator/admin rolesmanage queue, review drafts, watch computer-use sessions with plan-execute audit, Buku Kantor: browse/correct/delete Memori Aksara facts with provenance shown, consent management, device settingsno external access; LAN only
aksara-cli / MCPtechnical operators, agentseverything deterministic, scripted, auditednothing without a role-scoped key
ETNOSpublicverify records, see capability card, PII-free transparency feedno personal data, ever

The memory file manager Yose asked about is the dashboard's Buku Kantor view, not the wall panel: inspectable memory is a trust feature ("apa yang Aksara tahu tentang saya, perbaiki, hapus"), and it implements UU PDP correction/deletion rights concretely. The browser-use agent (Srikandi path) is driven from voice or dashboard; its progress shows in the statline, its full session log lives in the dashboard.

Voice identity pipeline (how voice ID coexists with Gemma ASR)

Speaker identification and speech recognition are different jobs and different models. Pipeline: mic array beam → VAD → two parallel branches: (1) a tiny speaker-embedding model (ECAPA-TDNN class, ~20 MB, milliseconds on CPU) produces a voice-fingerprint vector, matched by cosine similarity against enrolled staff embeddings stored in Memori Aksara's person table (consented, revocable); (2) Gemma 4 E4B transcribes WHAT was said. The runtime receives text + speaker attribution; memory writes carry that attribution and pass the consent gate. Models are stateless and store nothing; Memori Aksara (SQLite) is the only place anything persists. Honesty rule: voice ID is convenience-grade identity (greetings, preferences, case continuity), never approval-grade; approvals always require fingerprint or the WhatsApp liveness path. Anti-spoofing is therefore a UX problem, not a security boundary.

NVIDIA NIM note

NIM containers require NVIDIA hardware; the RK3588 (Mali GPU + NPU) cannot run them. Where NIM legitimately enters the stack: the data-residency cloud tier, because Deka LLM (Lintasarta) is built on NVIDIA AI Enterprise + NIM (v4.1 §15 already says this), and optionally a future Jetson-class Hub configuration (v4.1 roadmap). On the device: llama.cpp stays. So the answer is: NIM yes, but in the cloud tier where the NVIDIA silicon lives, never on the node.

Open questions

  • Tifa recording: who plays, and under what consent/credit (community mode question in miniature).
  • Punggung on left or right edge (right-handed stempel ergonomics suggest right).
  • Landscape vs portrait default per deployment package (MPP counter may want portrait queue view).
  • mmWave sensor placement inside the spine (verify LD2410 detection cone through PETG).
  • Greeting behavior defaults: spoken greeting on arrival may be per-institution config (some offices will love it, some will want halo-only).