SCHOLARLY ESSAY · AI GOVERNANCE, HUMAN AGENCY, AND CRITICAL TECHNOLOGY
FORTH and the Case for a Governed Personal–Executive AI
Why the next useful assistant must make invisible coordination work legible, retrieve reality, and preserve authorship under assistance.
Revised 7 August 2026 · Chicago author–date · 47-source bibliography · product documents treated as design evidence, not efficacy evidence
Autonomy is not the absence of assistance; it is the preservation of authorship under assistance.
The argument. Administrative capacity is practical power. This essay asks whether AI can distribute more of that capacity without reproducing the surveillance, classification, dependency, and invisible labor it promises to relieve.
CONTENTS
- The bottleneck after intelligence
- The hidden executive work of a life
- The work was never neutral
- Why a chatbot is not an assistant
- The jagged frontier changes the design problem
- Context is power: the dossier problem
- Personalization without a dossier
- One life, two domains
- From answer authority to action authority
- Agency is the product constraint
- Who gets to have an assistant?
- The strongest objection
- What FORTH is for
- Research note: source discipline
- References
The bottleneck after intelligence
In 1971 Herbert Simon made an observation that becomes more important, not less, when machines can generate competent prose, analysis, code, images, and plans on demand: “a wealth of information creates a poverty of attention” (Simon 1971, 40–41). His deeper point was architectural. An information-processing system earns its place in an organization only when it consumes less human attention than it saves. The useful system is therefore not the one that produces the greatest volume of information. It is the one that filters, condenses, buffers, and routes information so that scarce human judgment is spent where judgment matters.
Half a century later, generative artificial intelligence has lowered the cost of producing intellectual material while leaving Simon’s bottleneck intact. The age of AI has not eliminated the executive bottleneck; it has industrialized the production of material that can enter it. A language model can draft twelve versions of an email before a person could type one. It can propose itineraries, meeting agendas, recipes, research summaries, performance plans, or negotiation scripts. But the user must still determine which Daniel a message concerns; whether Thursday actually works; which calendar state is current; whether a private appointment may constrain a work commitment without being disclosed; whether a flight connection is robust door to door; whether a restaurant recommendation reflects tonight’s hours rather than last year’s webpage; whether a confident factual sentence is retrieved evidence or model inference; and whether “handle it” authorizes research, a draft, a reservation, a purchase, or an irreversible external act.
This residual work is easy to underestimate because it is distributed across dozens of small acts. It consists of remembering, disambiguating, searching, checking, comparing, sequencing, following up, correcting, and deciding what may safely cross a boundary. It is the work of maintaining coherence between a person’s intentions and the moving state of the world. Today the person commonly performs that work for the AI: retrieving the context, composing the prompt, checking the answer, moving information between systems, and executing the result. The ostensible assistant can therefore leave its principal serving as systems integrator.
FORTH is a proposed answer to that inversion. Its constitution does not define the product as a chatbot with calendar access. It specifies a context-aware personal and executive operating layer whose job is to infer low-consequence intent, retrieve authoritative facts, coordinate across live systems, rank options, prepare or execute authorized actions, preserve privacy boundaries, and close loops when the runtime actually permits it (FORTH Project 2026g). Its prime directive is revealing: infer harmless meaning aggressively; verify facts aggressively; become more conservative as consequences increase; finish the task when authority and capability are present. That is less a claim about a cleverer model than a claim about the architecture surrounding a model.
The updated design makes that architecture concrete. A compact always-on kernel carries the hot-path rules; ten versioned Sources divide stable context and operating doctrine into distinct roles; live tools retain ownership of changing truth. Executive Context is the first-use gate: eleven short, skippable prompts establish only the durable facts, preferences, boundaries, and support needs the user chooses to provide, while every blank remains unknown. People and relationship context, correspondence voice, and professional research routing are configured progressively rather than demanded as a biographical intake. Gmail owns current mail; Calendar owns recorded commitments; travel providers own current fares and schedules; authoritative web sources own changing public facts. A stored preference may rank the options returned by a live search, but it cannot make an old fare current or an inferred recipient verified (FORTH Project 2026c, 2026f, 2026g, 2026h, 2026j). The design is therefore not maximal memory. It is typed context with explicit ownership.
Its intended operating envelope is correspondingly broader than the earlier language of an "assistant with memory." Version 0.2 specifies one coordination system that can reconstruct broken dictation; resolve people, aliases, groups, and verified contact identities; calibrate correspondence to the user's own voice and to the relationship at hand; search and reason across current email and calendars; coordinate meetings; plan travel from door to door; route professional research through field-specific evidence standards; and create user-authorized scheduled checks for changing conditions. The current Life Logistics doctrine extends the same engine into meals, pantry and home-condition photographs, grocery consolidation, restaurants, household-safety triage, maintenance, warranty and recall research, local-service selection, credential verification, quote normalization, preventive maintenance, and project change control. The Human Performance source adds executive coaching, decision architecture, leadership systems, condition-aware support, and longitudinal adaptation that is expressly barred from becoming covert psychological profiling (FORTH Project 2026a, 2026d, 2026e, 2026g). These are specified capabilities, not a claim that every connector or permission is present in every session: the constitution requires runtime discovery and graceful degradation, and v0.2 excludes autonomous purchasing.
The apparent breadth is not feature sprawl if the unit of design is the completed human intention. The same grammar governs a mangled voice request, an email reply, a family trip, a contractor estimate, and an overloaded executive decision: reconstruct intent; retrieve the source that owns changing truth; combine it with only the relevant durable context; distinguish observation, retrieval, calculation, inference, recommendation, and unknown; decide what can be prepared or acted on under the user's authority; execute at most the authorized mutation; reconcile the result; and report the narrowest state the evidence establishes. FORTH's object is therefore not a collection of apps. It is continuity across the seams where human attention is usually spent joining apps, facts, relationships, and obligations together (FORTH Project 2026g).
The case for such an architecture rests on a simple thesis. Modern AI has made generic cognitive output abundant. What remains scarce is contextualized, accountable follow-through: attention, prospective memory, coordination, and judgment under changing constraints. If AI is to become an assistant in the ordinary meaning of the word—something that reduces the work required to bring an intention into contact with reality—then intelligence alone is insufficient. It needs a governed context-and-action layer. FORTH is one concrete design hypothesis for that missing layer.
That is also the appropriate sense in which it is necessary. No empirical literature proves that this particular product must exist, and a serious argument should not pretend otherwise. The necessity claim is categorical and conditional: once an AI is asked to operate across the real systems of a human life, some mechanism must solve context, provenance, authority, privacy, and completion. A future foundation model may become dramatically more capable, yet greater model intelligence cannot by itself confer today’s flight status, decide which database owns a changing fact, authorize a payment, establish whether a private fact is appropriate to disclose to a colleague, or prove that an external mutation actually occurred. Those are not residual IQ problems. They are problems of system design and governance.
The hidden executive work of a life
The unit of work in everyday life is rarely the isolated “task.” Allison Daminger’s qualitative study of household cognitive labor offers a more revealing grammar. Across seventy interviews with thirty-five couples, she distinguished four activities: anticipating needs, identifying options, deciding among them, and monitoring results. In that sample, women performed more cognitive labor overall and more of the anticipation and monitoring in particular (Daminger 2019). The study concerns heterosexual couples and a specific qualitative sample, not executive work at large; it should not be converted into a population-wide productivity estimate. Its conceptual value is nevertheless substantial. The same sequence recurs in professional and personal administration. “Book dinner” may require noticing that dinner is needed, locating feasible restaurants, reconciling distance and timing with a preceding event, selecting among trade-offs, booking within the user’s authority rules, placing the result on a calendar, and later detecting a schedule change that invalidates the plan. The visible click is the end of a cognitive chain.
Psychology supplies a complementary account. People routinely move prospective-memory demands into calendars, diaries, objects, alarms, and smartphone reminders. Gilbert and colleagues’ review calls this “intention offloading” and finds that external reminders can be highly effective, while also showing that people’s decisions about when to offload are shaped by metacognition and systematic biases (Gilbert et al. 2023). Risko and Gilbert place this behavior within the broader category of cognitive offloading: changing the environment through action so that a cognitive demand is reduced (Risko and Gilbert 2016). The point is not that memory should be outsourced indiscriminately. It is that competent human action has always depended on intelligently arranged external supports.
The qualification in that last sentence is essential. In three experiments, Grinschgl, Papenmeier, and Meyerhoff found a characteristic offloading trade-off: externalizing information could improve immediate task performance while diminishing subsequent memory for what had been offloaded (Grinschgl, Papenmeier, and Meyerhoff 2021). The implication for an assistant is not “remember everything for me.” It is to externalize the right burden while preserving enough understanding and recoverability for the person to remain an effective principal.
Distributed-cognition research pushes the argument further. In Cognition in the Wild, Edwin Hutchins showed that navigation is accomplished through a system of people, representations, instruments, procedures, and transformations rather than by treating cognition as something sealed inside an individual skull (Hutchins 1995). Lucy Suchman’s work on situated action likewise challenged machine conceptions in which a plan can fully determine action in advance; real conduct is produced in ongoing interaction with local circumstances (Suchman 1987). Neither work is experimental evidence that FORTH will succeed. They are better used for what they actually establish: a change in the unit of analysis. The question is not simply whether the model is intelligent. It is whether the human–model–tool system preserves the information and relationships required for intelligent action in context.
This makes prompt burden more than a usability annoyance. If the user has to restate the relevant people, commitments, preferences, previous corrections, travel constraints, privacy rules, and live state every time, the system repeatedly charges attention for context it ought to manage. Task switching compounds the problem: laboratory work by Rubinstein, Meyer, and Evans found measurable switching costs that varied with task complexity and advance preparation (Rubinstein, Meyer, and Evans 2001). That literature does not justify a simplistic claim that “multitasking destroys productivity.” It supports the narrower observation that control itself has costs. A useful assistant should therefore reduce unnecessary control operations without hiding consequential choices.
FORTH’s internal design documents read as an attempt to make that principle operational. Email and calendar are not conceived as separate “features” but as a coordination substrate: reconstruct enough thread state to know whether something is a proposal, an acceptance, or silence; resolve the actual recipient before acting; refresh stale calendar state before making a commitment; let a private event block availability without disclosing its reason (FORTH Project 2026b). Travel is treated not as a list of flights but as a door-to-door optimization problem whose evidence is tagged by source, owner, time, scope, and status; the plan is stress-tested against calendar anchors and recovery options (FORTH Project 2026i). Food and dining similarly become decision problems in which live hours, dietary constraints, geography, household state, price, and friction may matter more than prestige (FORTH Project 2026e). Across domains, the common object is not content generation. It is the removal of coordination work.
Simon gives us the test. A tool that saves thirty seconds of drafting while creating two minutes of checking, copying, disambiguating, and recovering context has not conserved attention. It has displaced the clerical burden into a new interface.
The work was never neutral
Simon names scarcity. Feminist scholarship asks a question that scarcity language can hide: whose attention has historically been spent keeping other people's lives coherent? Once the question is posed, the background changes. Anticipating meals, remembering appointments, coordinating children and elders, maintaining the home, noticing supplies, arranging care, managing social obligations, and monitoring whether any of it actually happened are not newly invented "AI workflows." They belong to the long history of social and reproductive labor. Daminger's study makes one cognitive component of that labor analytically visible; its gender pattern also warns against treating the work as if it fell randomly across households (Daminger 2019).
Race complicates any account that stops at "women." Evelyn Nakano Glenn's historical analysis of U.S. reproductive labor traced continuities between domestic service and later institutional service work, examining how race and gender were constructed together in a division of labor that placed African American, Mexican American, Japanese American, and other racialized women disproportionately in particular forms of paid reproductive work (Glenn 1992). Glenn's article is historical sociology, not a demographic estimate for 2026. Its importance here is structural: one person's relief from maintenance and care work has often been produced by another person's labor, and that transfer has been organized through race, gender, class, immigration, and occupational hierarchy. A technology that proposes to "remove life admin" enters that history whether its designers acknowledge the history or not.
Black feminist thought supplies the more demanding analytic rule. The Combahee River Collective's 1977 statement insisted on integrated analysis because racial, sexual, heterosexual, and class oppression were interlocking in lived experience (Combahee River Collective 1977). Kimberlé Crenshaw later demonstrated how single-axis legal and political frameworks could make Black women's discrimination analytically disappear when race and sex were treated as separate, mutually exclusive channels (Crenshaw 1989). Patricia Hill Collins identified self-definition and self-valuation, the interlocking character of oppression, and Black women's culture as central themes in Black feminist thought (Collins 1986). None of these texts is a product specification, and none should be conscripted into a retrospective endorsement of personal AI. They do something more valuable: they expose the inadequacy of a universal user whose burdens, risks, and access are presumed to be the same.
That lesson changes what counts as a successful assistant. A system optimized for an archetypal executive may save time for people who already possess money, institutional fluency, staff, stable broadband, predictable schedules, and low exposure to bureaucratic scrutiny. The same interface can impose more work on someone navigating disability, care obligations, unstable hours, unfamiliar institutional language, or repeated demands to prove eligibility and identity. Research on administrative burden is especially instructive. Victor Ray, Pamela Herd, and Donald Moynihan argue that learning, compliance, and psychological costs imposed by public administration can function as racialized mechanisms of inequality even when the rules that carry them appear facially neutral (Ray, Herd, and Moynihan 2023). In a qualitative study drawing a stratified sample of sixty-one Black, Latinx, and White women from the American Voices Project, Theresa Rocha Beardall, Collin Mueller, and Tony Cheng found that social location shaped how participants experienced and navigated burdens in income assistance, health care, and housing (Rocha Beardall, Mueller, and Cheng 2024). That study concerns encounters with the administrative state; it does not establish that FORTH will reduce those burdens or that its findings generalize to every institutional setting. It establishes the point a productivity metric can miss: friction can be distributional.
This is where the argument for FORTH becomes more consequential than "save busy people time." Administrative capacity is a form of practical power. The ability to find the governing rule, assemble the required evidence, remember the deadline, identify the correct professional, compare the options, phrase the communication, protect a boundary, and follow the matter through completion changes what a person can actually accomplish. Some people buy that capacity through executive assistants, household managers, attorneys, travel advisors, accountants, care workers, or domestic labor; others supply it personally, draw it from family, or go without. A governed AI cannot erase the institutions or inequalities that produce those differences. It can, in principle, make a portion of high-quality coordination capacity cheaper to reproduce and easier to carry across domains. That is a capability claim worth testing, not a social-justice accomplishment to be announced in advance.
The feminist implication is therefore double. First, the labor should be made visible enough to design for it: anticipation and monitoring are part of the task, not residue left to the user after an impressive answer. Second, automation must not make the labor socially invisible again. Catherine D'Ignazio and Lauren Klein's Data Feminism treats data work as a field of power and explicitly calls attention to the human labor hidden inside ostensibly automated systems (D'Ignazio and Klein 2020). FORTH should be judged by the same standard. If the system's apparent autonomy depends on users repeatedly correcting context, family members maintaining its data, low-paid workers cleaning its training or retrieval pipeline, or human assistants silently repairing its errors, then the labor has not disappeared. It has changed location.
Nor is all friction waste. Joan Tronto's political ethic of care treats care as a practice entangled with responsibility, competence, responsiveness, and power rather than as a naturally feminine disposition (Tronto 1993). Applied cautiously to AI, that distinction prevents a serious category error. Searching three pharmacies, comparing contractor scopes, reconciling calendars, or remembering a refill can be burdensome coordination that a user may gladly delegate. Listening to a frightened child, making amends to a spouse, deciding how to support a dying friend, or giving difficult feedback to a colleague contains relational and moral work that cannot be reduced to the efficiency of its logistics. An assistant may prepare the scaffolding; it should not tell itself that the relationship has been "handled." The relevant boundary is not personal versus professional. It is instrumental burden versus authored responsibility.
Why a chatbot is not an assistant
Consider a realistic dictation: “Push Daniel to Thursday, make sure tomorrow works, and find somewhere actually good for dinner after I land—not a production.” A general-purpose chatbot may respond fluently. But the language of the request underspecifies almost everything an assistant needs to know. Which Daniel? What existing commitment is being pushed? Does “tomorrow works” refer to the itinerary, the calendar, the briefing, or all three? Where is the traveler landing, at what time, and with what plausible ground delay? Does an undisclosed personal commitment eliminate an apparently free Thursday slot? What does “actually good” mean for this user, on this trip, with this party, at that hour? Does “push” authorize sending a message or only preparing it?
A human executive assistant handles such sentences by combining relationship knowledge, situational cues, current records, tacit preferences, and an implicit model of the principal’s authority regime. The assistant makes many low-risk inferences silently and escalates the few ambiguities that could change a commitment. FORTH formalizes that pattern. Its constitution separates retrieved facts, durable user facts, calculations, inferences, recommendations, and unknowns; it forbids presenting inference as retrieval. It also separates what the system is specified to do from what is configured, tested, and actually passing. A source or tool described in a document is not itself production evidence that the capability exists (FORTH Project 2026g).
The new onboarding Sources make an important part of that pattern inspectable. Executive Context records stable user-owned facts and preferences while keeping live account state out of the profile. The People source treats a group roster as a dated assertion, not an eternal audience, and refuses to upgrade a guessed email address through repetition or name similarity. The Voice source makes an equally subtle separation: approved samples can teach cadence, warmth, compression, and register, but the names, dates, promises, attachment claims, and shared history inside those samples do not become facts in a new message. The Research Field Pack stores authority-routing rules and freshness thresholds, not current answers (FORTH Project 2026c, 2026f, 2026h, 2026j). These distinctions matter because context is useful only when the system also knows what the context is evidence for.
That last distinction is unusually important. AI products invite a category error between linguistic competence and operational competence. A model can describe how it would check a calendar without possessing the calendar; describe a booking without making one; compose a plausible restaurant detail without consulting a current owner of that fact; or state that it will remember a preference when no durable writeback has been verified. Fluency obscures the gap because a simulation of completion is expressed in the same medium as completion itself: words.
A governed assistant has to make state legible. What is known? From where? How fresh is it? What was calculated? What was inferred? Which source owns the changing truth? What action is proposed? Has the user delegated it? Was it actually executed? Can its success be verified? FORTH’s truth typing, live-source routing, writeback verification, and action-authority rules are answers to these questions (FORTH Project 2026g). They are not decorative safety features added after intelligence. They are what allow intelligence to become operational without forcing the user to audit every sentence and click.
This also explains why “better prompting” is an inadequate product philosophy. Prompting can improve an answer to a bounded question. It cannot reasonably require a person to reinstantiate the state of a life. Suchman’s situated-action critique matters here: action is revised as circumstances unfold, so the system must remain coupled to the circumstances that can invalidate its plan (Suchman 1987). Hutchins matters for the same reason: the relevant cognitive system includes external representations and tools (Hutchins 1995). The assistant’s task is therefore not to hold an ever larger private monologue about the user. It is to maintain disciplined contact between intent and the appropriate external state.
The jagged frontier changes the design problem
The case for an operating layer would be weak if generative AI itself offered little leverage. The strongest evidence says otherwise—but only when stated with discipline.
In a preregistered online experiment involving 453 college-educated professionals completing occupation-specific writing tasks, Shakked Noy and Whitney Zhang found that access to ChatGPT reduced completion time by about 40 percent and increased evaluator-rated output quality by about 18 percent (Noy and Zhang 2023). The result is important precisely because it is bounded: it demonstrates substantial gains on a class of professional writing tasks, not a forty-percent productivity law for knowledge work.
Evidence from a deployed workplace is similarly consequential. Erik Brynjolfsson, Danielle Li, and Lindsey Raymond studied the staggered introduction of an AI assistant to 5,172 customer-support agents. Access increased issues resolved per hour by 15 percent on average, with the largest gains among less-experienced and lower-skilled workers; the most experienced workers saw small speed gains and small declines in conversation quality (Brynjolfsson, Li, and Raymond 2025). Again, the setting matters. This was one support environment with a particular tool and workflow. The study shows that a well-positioned generative system can change real operational performance; it does not show that a generic model improves every worker or task.
The more decisive evidence for FORTH’s architecture comes from where assistance fails. Dell’Acqua and colleagues conducted a preregistered randomized experiment with 758 consultants. Across eighteen tasks chosen to sit inside the model’s capability frontier, AI users completed 12.2 percent more tasks and did so 25.1 percent faster on average, with average response quality 32 percent higher. On a deliberately selected complex task outside that frontier, however, pooled correctness in the AI conditions was about nineteen percentage points lower than in the control condition (Dell’Acqua et al. 2026). The experiment does not map the full boundary of modern AI—its “outside” finding comes from a single task, and capabilities evolve. Its importance lies elsewhere: tasks that look similarly difficult to people can lie on opposite sides of an AI system’s competence boundary. The frontier is jagged.
A large secondary synthesis makes the warning harder to dismiss as an isolated result. Michelle Vaccaro, Abdullah Almaatouq, and Thomas Malone reviewed seventy-four papers comprising 106 experiments and 370 effect sizes in which human-only, AI-only, and combined human–AI performance could be compared. Human–AI systems outperformed humans alone on average, but they underperformed the better of the human-only or AI-only baselines on average; task type materially affected the result (Vaccaro, Almaatouq, and Malone 2024). A human plus an AI is therefore not automatically a superior decision system.
These findings shift the engineering target. If model capability were smooth and reliably self-evident, an assistant could simply hand every problem to the model and trust the answer. The evidence suggests the opposite. A useful system must route tasks, detect when live evidence is required, preserve the difference between generated inference and retrieved fact, and become more conservative as error costs rise. NIST’s Generative Artificial Intelligence Profile similarly treats confabulation, data privacy, human–AI configuration, and source verification as risk-management concerns, rather than assuming that fluent output establishes reliability (NIST 2024). NIST is official guidance, not efficacy evidence for FORTH. Its relevance is normative and architectural: reliable deployment requires controls beyond the generative model.
This is why FORTH’s most important feature may be the least glamorous one: it refuses to collapse epistemic states. A restaurant’s current hours, a flight’s status, a contact’s identity, a user’s remembered preference, a calculated transfer time, and a model’s recommendation are different kinds of claims. Treating them as interchangeable is convenient until the assistant acts. Once an AI moves from saying to doing, provenance becomes part of competence.
Context is power: the dossier problem
The argument becomes harder when the thing being made operational is a person. A personal–executive system becomes useful by accumulating enough context to stop asking the user to restate a life. The same accumulation can produce a remarkably convenient apparatus of surveillance. The difference cannot be settled by saying that the data are "personalized" or that the user once clicked consent. The governing questions are who supplied the claim, what the claim is evidence for, what may be inferred from it, how long it persists, who may receive it, what purpose it may serve, whether it can be corrected, and what action it can authorize.
Feminist data scholarship makes power, rather than volume, the proper unit of analysis. D'Ignazio and Klein argue that data science is a form of power and that context, invisible labor, plural knowledge, and the politics of classification have to be examined rather than assumed away (D'Ignazio and Klein 2020). Black studies makes the surveillance problem still less abstract. Simone Browne's Dark Matters places modern surveillance in a longer history of racializing surveillance, tracing connections among the surveillance of Black life, identification, bordering, biometrics, and social sorting (Browne 2015). Ruha Benjamin's Race After Technology examines how technological systems can reproduce and deepen racial hierarchy while appearing innovative or neutral (Benjamin 2019). These works do not show that a context-aware assistant is inherently racist or surveillant. They make a narrower but unavoidable point: technical mediation does not dissolve the social histories and power relations embedded in classification, observation, and control.
Primary empirical research supplies a concrete warning about universal performance claims. Joy Buolamwini and Timnit Gebru evaluated three commercial gender-classification systems and found large intersectional disparities: in their benchmark, darker-skinned women were the most frequently misclassified subgroup, with error rates reaching 34.7 percent, while the maximum reported error rate for lighter-skinned men was 0.8 percent (Buolamwini and Gebru 2018). Facial gender classification is not personal-assistant inference, and the study should not be treated as if it benchmarked language models. Its lesson is methodological. An apparently competent system can have sharply unequal failure rates that aggregate accuracy conceals. "It works" is not an adequate evaluation claim until one asks for whom, under which conditions, and at what cost when it fails?
FORTH's architecture contains several unusually relevant answers, though at present they are design commitments rather than proof. Its first-use Executive Context is user-authored: unanswered fields remain unknown rather than becoming invitations to infer identity or preference from demographic priors. Its People source separates a verified identifier from a guessed one. Its voice model may learn register from approved samples but may not treat the facts inside an old sample as current history. Its Human Performance model expressly forbids building durable psychiatric, personality, political, sexual, medical, or moral profiles from response latency, dictation errors, inbox hygiene, purchases, browsing, third-party labels, or other operational traces. It separates temporary state, situation-specific pattern, explicit preference, and hypothesis; it makes current correction outrank historical pattern; and it prohibits mining connected systems solely to enrich a psychological model (FORTH Project 2026c, 2026d, 2026f, 2026j).
Here Collins's emphasis on Black women's self-definition provides a particularly exacting normative lens. Collins was theorizing Black feminist knowledge and resistance to objectifying definitions, not database schemas (Collins 1986). The translation should therefore remain explicit and limited: a system serving a person should not acquire unilateral authority to define that person's identity from behavioral residue. FORTH's maxim that a current user correction outranks an inferred pattern echoes a principle of epistemic self-authorship; it does not fulfill Black feminist politics by technical analogy. The distinction matters because the most polished form of algorithmic paternalism is often the system that says, in effect, I know you better than you know yourself.
Catherine Knight Steele's Digital Black Feminism supplies a complementary corrective to a risk-only account. Her history of Black feminist technoculture centers Black women not simply as populations acted upon by technology but as skilled users, makers, entrepreneurs, intellectuals, and agents whose practices have shaped digital culture (Steele 2021). That matters for FORTH's design imagination. The objective cannot be to construct a more benevolent system about marginalized users. It is to construct a system in which users remain authors of the context, purposes, corrections, boundaries, and forms of assistance that govern them. User agency is not a safety feature downstream of personalization. It is the legitimacy condition for personalization.
This is why FORTH's distinction among user-authored facts, system inferences, live external truth, and action authority is more than tidy information architecture. Each category answers a different power question. User-authored context says what the person has chosen to establish. Inference says what the system is tentatively proposing. Live retrieval says what an external source currently records. Authorization says what the assistant is permitted to change. Collapsing any pair creates a recognizable abuse: inference becomes identity; memory becomes truth; access becomes disclosure; recommendation becomes consent; technical capability becomes permission.
The architecture still has an unresolved weakness that critical theory makes impossible to hide. Behavioral rules inside a project are not the same thing as a technically enforced privacy plane, and a well-written anti-profiling policy is not evidence that all users will experience the system as contestable or safe. Participatory testing across people whose risk profiles differ is not currently established by the Source set. Nor does individual consent cure every structural problem created by concentrated data or unequal access. The honest claim is therefore architectural and falsifiable: FORTH is designed to obtain the benefits of continuity without making silent classification the price of convenience. Whether it succeeds has to be measured in real use.
Personalization without a dossier
An assistant that never learns its principal remains expensive to use. An assistant that learns indiscriminately becomes something else: a profiling system with unusually intimate access. The difficult design problem is not whether to personalize. It is what personalization is allowed to mean.
FORTH’s updated Human Performance source makes this problem explicit through an Adaptive Personal Operating Model. Its unit is not a personality type but a conditional relationship: person × current state × situation × task × assistance → observed outcome. The system is allowed to learn that one visible next action has repeatedly helped in a certain class of overloaded tasks, or that the user explicitly prefers adversarial analysis for identity-relevant career decisions. It is not allowed to turn messy dictation, response latency, inbox behavior, purchasing patterns, a third party’s description, or a single difficult week into a durable diagnosis or trait. The source separates current instructions, explicit standing preferences, outcome feedback, operational outcomes, repeated low-sensitivity behavior, and hypotheses by evidentiary weight; it also separates four outcomes that personalization systems can easily conflate: what the user says they prefer, what feels helpful, what actually changes the target behavior or decision, and what unintended costs follow (FORTH Project 2026d).
There is empirical reason to prefer this conditional approach, although the evidence should not be stretched into product validation. Emorie Beck and Joshua Jackson used 5,971 intensive longitudinal assessments from 104 college-age adults to build person-specific predictions of loneliness, procrastination, and studying. Both person and situation variables contributed, while the most important predictive features varied substantially across individuals (Beck and Jackson 2022). This supports a narrow proposition: person-specific prediction can contain information obscured by one-size-fits-all models. It does not show that conversational traces reveal hidden personality, that the same predictors generalize to executives or clinical outcomes, or that a personalized intervention will cause improvement.
Personal context can also make a model worse in a peculiarly seductive way: it can make the model feel more attuned while weakening its independence. Shomik Jain and colleagues studied two weeks of interaction context from thirty-eight users and evaluated agreement and perspective sycophancy across several language models. Agreement sycophancy tended to increase when context was supplied, with effects varying materially by model and context type; perspective mirroring increased when models could infer users’ viewpoints (Jain et al. 2026). The study is bounded—thirty-eight users, a limited observation period, particular models and operationalizations—and it does not establish that all memory or personalization produces sycophancy. It does establish something a serious personal assistant must test rather than assume: more knowledge about the user can change the epistemic behavior of the model, not just its convenience.
FORTH’s answer is an epistemic firewall between a reality layer and a personalization layer. Retrieval, evidence quality, counterevidence, uncertainty, causal discipline, privacy, safety, and authorization are supposed to be decided without allowing the user profile to rewrite them. Personalization then changes presentation, priority among defensible options, pacing, friction, and selection among interventions that survive the reality layer. A current instruction overrides a historical preference. A learned pattern is narrowed or retired when counterevidence appears. Sensitive psychological hypotheses remain task-local by default. Connected email, calendar, files, or browsing are not to be mined simply to enrich a psychological portrait. If personalization materially changes a consequential recommendation, the basis should become proportionately visible without exposing private chain-of-thought (FORTH Project 2026d).
This distinction refines the older language of “memory.” The best personal assistant should not know the greatest possible number of things about its user. It should retain the smallest set of well-governed facts, preferences, relationships, and response patterns that predictably remove repeated explanation while remaining correctable and scoped. The relevant measure of learning is not intimacy. It is reduced coordination cost without reduced epistemic independence. The assistant should learn how to help, not acquire the authority to decide who the person is.
One life, two domains
The phrase personal plus executive can sound like a bundling strategy: email beside dinner reservations, meetings beside travel, coaching beside groceries. The stronger interpretation is architectural. A human being lives one life under multiple disclosure regimes.
The workday and the personal day compete for the same hours, body, attention, relationships, and geography. A medical appointment may make an executive unavailable. A child-care constraint may determine which flight is viable. A business trip may change household logistics. A dinner may be both social and professional. The assistant therefore needs some cross-domain context to avoid making locally rational and globally impossible plans. Yet the fact that a personal constraint is relevant to work does not make its content appropriate for work disclosure.
Helen Nissenbaum’s theory of contextual integrity provides a useful vocabulary. Privacy is not exhausted by secrecy or by a binary choice between “shared” and “private”; it concerns whether information flows are appropriate to the norms of a context—who sends what information about whom to whom, under what conditions (Nissenbaum 2004). For an integrated assistant, this yields a demanding design principle: constraints may sometimes cross a boundary when secrets should not. The work calendar may need to know “unavailable from 3:00 to 4:30,” while a colleague has no need to know why.
FORTH’s constitution explicitly adopts this pattern. It describes Personal and Executive as operating domains within one coordinated system and allows cross-domain context to be used privately while imposing stricter rules on disclosure (FORTH Project 2026g). Just as importantly, it admits the limit of the current design: in a project-level implementation, that separation is a behavioral policy, not cryptographic isolation; a standalone product would require technical enforcement appropriate to its threat model. This is not a footnote to the product thesis. It is evidence of whether the thesis is intellectually serious. Integration without boundary control is surveillance. Boundary control without integration leaves the user doing the reconciliation.
Yet “Personal” and “Executive” cannot themselves be the complete privacy model. Contextual integrity asks more than which side of a two-domain line a datum occupies. Nissenbaum’s later formulation makes the relevant parameters explicit: sender, recipient, subject, information type, and transmission principle; she also argues that downstream use may need explicit treatment (Nissenbaum 2019). A private fact can therefore be mishandled inside an integrated system even if it is never visibly disclosed across the Personal/Executive boundary—for example, if it is silently repurposed to influence a decision the user would not expect it to inform. The two domains are a useful first partition, not a sufficient theory of permissible information flow. A mature FORTH would need purpose- and flow-sensitive controls alongside domain separation.
The design problem is thus not to erase the border between personal and executive life. It is to let relevant constraints cross without letting secrets follow them. Integration is not data centralization. It is governed contextual passage.
From answer authority to action authority
An assistant becomes qualitatively different when it can mutate the world. A recommendation can be ignored; a sent email creates a social fact. A search is reversible; a purchase, cancellation, deletion, or public disclosure may not be. The proper question is therefore not simply, “Can the AI do this?” but “By what authority may it do this now?”
FORTH’s design distinguishes research and preparation from external mutation. It expects the system to infer and proceed more freely when consequences are low, while requiring stronger certainty or confirmation as recipient, commitment, disclosure, cost, irreversibility, or sensitivity increase (FORTH Project 2026g). The email/calendar playbook makes recipient resolution a safety gate: a plausible wrong recipient is still wrong (FORTH Project 2026b). The travel playbook distinguishes a decision-ready or booking-surface handoff from an actual booking (FORTH Project 2026i). These distinctions are mundane in the best sense. They encode what competent human assistants already know: authority is granular.
Tool use introduces a second authority problem that is easy to miss. Once an assistant reads email, webpages, attachments, or calendar descriptions, it is ingesting text written by people who are not the principal. Kai Greshake and colleagues demonstrated indirect prompt-injection attacks in which adversarial instructions embedded in data retrieved by LLM-integrated applications could manipulate model behavior and tool use (Greshake et al. 2023). The precise attack surface has evolved since that work, but the category does not vanish when a model becomes more capable: data and authority must remain distinguishable. FORTH therefore treats retrieved content as evidence, never as permission. A sentence inside an email can supply a date to a workflow; it cannot grant itself the right to send another email, retrieve unrelated private data, change permissions, or override the user’s disclosure rules (FORTH Project 2026g). This is not conventional “prompt hygiene.” It is a control-plane requirement for any assistant that reads untrusted language and can affect external systems.
The updated email/calendar design extends the same principle to transactional integrity. Immediately before a time-sensitive send or calendar write, decisive mutable state is refreshed. If a write times out or returns an ambiguous result, the system is instructed to reconcile against the provider’s canonical state before retrying rather than risk a duplicate. Partial success remains partial success; provider acceptance of a request is not inflated into proof of downstream delivery, attendance, or completion (FORTH Project 2026b). In other words, competent action is not the instant at which the model decides what to do. It is a sequence: refresh, authorize, mutate once, reconcile, and report the narrowest state the evidence actually establishes.
Proactivity makes the authority problem temporal. A recurring task created today may run tomorrow in a changed world. FORTH’s scheduled-work doctrine therefore gives an automation only the standing authority explicitly granted at creation, requires it to recheck stale observations before interrupting the user, favors meaningful change detection over repetitive summaries, and withholds fresh destructive, sensitive, reputational, or financial action authority unless the relevant confirmation condition is satisfied (FORTH Project 2026g). A useful chief-of-staff heartbeat should lower interruption load; it should not become a standing power of attorney.
Friction, on this account, is not inherently an interface failure. In an experiment with 199 participants, Buçinca, Malaya, and Gajos found that interfaces designed to force more deliberate engagement reduced overreliance on incorrect AI advice relative to simpler assistance, although participants also experienced the more effective interventions as more burdensome and preferred them less (Buçinca, Malaya, and Gajos 2021). The study used one low-stakes decision task, so it cannot validate FORTH’s confirmation thresholds. It does establish the design tension: the least effortful interface is not always the one that best preserves judgment. The target is located friction—near consequential uncertainty, not everywhere.
The history of automation gives reason to resist both extremes—constant human micromanagement and indiscriminate autonomy. Lisanne Bainbridge’s classic “Ironies of Automation” observed that automation can leave human operators with the hardest monitoring and exception-handling work, sometimes while eroding the very practice that keeps those skills sharp (Bainbridge 1983). Her paper predates language models by four decades and should be treated as a human-factors analogy, not direct evidence about them. But the structural warning travels well. If an AI executes routine cases and summons the human only at the boundary, the system must preserve situation awareness, meaningful control, and the ability to inspect what happened. Otherwise “human in the loop” can become a ceremonial phrase for a person who bears responsibility without usable agency.
Later evidence gives the warning empirical weight while retaining the domain caveat. A meta-analysis of eighteen experiments on levels and stages of automation found a trade-off: higher automation tended to improve routine performance and reduce workload while degrading situation awareness and performance when automation failed (Onnasch et al. 2014). These studies largely concern supervisory-control settings rather than language-model assistants, so the transfer is structural rather than numerical. They establish why failure recovery belongs in the design brief; they do not provide an expected effect size for FORTH.
FORTH’s answer should therefore be judged by a demanding standard: does it relocate judgment rather than anesthetize it? The system should absorb clerical cognition—lookup, comparison, reconciliation, reminder, drafting, monitoring—while surfacing choices in which values, accountability, ambiguity, or irreversibility actually require the principal. That is an argument for selective automation, not maximal automation.
Agency is the product constraint
The deepest danger in personal AI is not a spectacular error. It is a gradual inversion of means and ends in which the person becomes dependent on the system that was supposed to enlarge their capacity.
Ivan Illich’s Tools for Conviviality supplies a useful normative provocation: good tools enlarge the scope of autonomous and creative action rather than reorganizing their users into dependents of an industrial system (Illich 1973). Amartya Sen’s capability approach supplies another: development should be assessed in terms of substantive freedom—what people are genuinely able to do and be—rather than by resources alone (Sen 1999). Neither author could validate an AI assistant they never saw. Their value is as tests of purpose. The right outcome is not the maximum number of decisions made by the machine. It is greater effective agency for the person.
The FORTH coaching source states this unusually plainly: the user should become more capable over time, not more dependent on the assistant (FORTH Project 2026d). Version 0.5 strengthens that principle by making “optimize the user’s objective, not product engagement” an operating rule and by treating failed personalization as correction data rather than resistance. Its coaching model is constrained by an evidence hierarchy, prohibitions on covert diagnosis and manipulation, and the rule that complex psychological models should remain quiet unless they materially improve the smallest useful intervention. That restraint matters because the evidence for AI coaching is still immature. A recent systematic review by Jonathan Passmore, Bergsveinn Olafsson, and David Tee found a promising but narrow and heterogeneous research base; it does not warrant treating AI coaching as interchangeable with expert human coaching, still less with clinical care (Passmore, Olafsson, and Tee 2026). The correct product stance is therefore bounded ambition.
The stronger principle is authored delegation: automate means while preserving the user’s authorship of ends. The person should retain control over purposes, recipients, disclosures, and consequential commitments even when the system performs much of the work between intention and outcome. Autonomy is not the absence of assistance; it is the preservation of authorship under assistance. Ryan and Deci’s account of self-determination explicitly separates autonomy—self-endorsed regulation—from independence, allowing that dependence on others can itself be autonomously chosen (Ryan and Deci 2006). The same distinction matters for tools. The goal is not to make the person perform every operation unaided. It is to ensure that delegated operations remain intelligibly connected to purposes the person can inspect, revise, or revoke. Sen’s language of substantive freedom and Illich’s concern for tool-shaped dependence converge on that test without pretending to provide a product benchmark (Illich 1973; Sen 1999).
Agency also changes how memory should be understood. Personalization is not the silent accumulation of everything the system can infer. A preference observed once may be situational. A correction may supersede an old rule. A sensitive fact may be useful now but inappropriate to retain. FORTH’s context contract therefore distinguishes durable explicit fact, durable preference, episodic context, observed pattern, and hypothesis, with different authority and lifecycle rules; the new Executive Context form additionally makes blank fields explicitly unknown rather than invitations to demographic inference (FORTH Project 2026c, 2026g). Memory should have provenance, scope, contestability, and a verified writeback path. The user should be able to correct what the system relies on, and the system should claim persistence only after persistence actually occurs. Otherwise personalization becomes epistemic capture: the system’s model of the person slowly acquires authority over the person.
The aspiration of FORTH is stronger: repeated use should reduce the amount of procedural labor the user performs while preserving—or increasing—the user’s ability to make consequential choices. The product should learn enough that the human no longer has to explain the mechanics of their life, but never so aggressively that yesterday’s inference becomes tomorrow’s identity.
Who gets to have an assistant?
There is a distributional question hidden inside the product category. High-quality administrative capacity has traditionally been expensive. Executives can hire people to screen information, maintain calendars, prepare travel, track commitments, resolve logistics, and protect their attention. Affluent households can purchase cleaning, food preparation, childcare, bookkeeping, maintenance management, and other forms of reproductive labor. Many people do some or all of the same coordination themselves, often alongside paid work and care obligations. The difference is not simply convenience. It changes the amount of cognitive slack available for judgment, recovery, relationships, and opportunity.
This is the strongest sense in which a governed personal–executive AI could become critical technology. The term should not be confused with a formal government designation of critical infrastructure. The claim is functional: if administrative capacity helps determine whether people can convert intentions and rights into completed outcomes, then making reliable administrative capacity more widely available can enlarge practical capability. Sen's framework is useful precisely because it refuses to equate possession of a resource with what a person is actually able to do with it (Sen 1999). A calendar, a health benefit, a legal right, a plane ticket, a pantry, a maintenance budget, or an inbox is not self-executing. Each can impose a conversion problem between nominal resource and usable outcome.
FORTH's breadth matters here. A narrow executive email agent primarily accelerates one professional workflow. A system that can also understand a household constraint, repair mangled dictation, research an unfamiliar professional domain, recognize that a personal obligation blocks a work commitment without disclosing why, turn a home symptom into a safe provider search, normalize two opaque contractor quotes, plan food around what is actually visible in the refrigerator, adapt task structure to a disclosed disability or low-capacity state, and return to a future condition only when the user has authorized a scheduled check is addressing a different object: the conversion of fragmented information into usable personal agency (FORTH Project 2026d, 2026e, 2026g). The capability is most important in the seams, because seams are where people with abundant support already employ humans to maintain continuity.
But democratization does not follow from software availability. A system can just as easily amplify existing advantage. If the best models, connectors, privacy protections, and support are available primarily to already powerful users, then automated administrative capacity becomes another compounding asset. If onboarding assumes professional vocabulary, stable institutions, one kind of family, standard speech, or predictable time, it can transfer prompt burden onto the people it purports to help. If errors are cheap for an executive choosing dinner but costly for a person navigating benefits, housing, immigration, medical care, employment, or debt, average task success is an ethically weak metric. The administrative-burden literature makes this asymmetry visible without proving that AI can solve it (Ray, Herd, and Moynihan 2023; Rocha Beardall, Mueller, and Cheng 2024).
An intersectional evaluation would therefore ask more than whether the mean user saves minutes. It would examine whether setup burden, clarification burden, decisive-fact errors, stale-state errors, privacy leaks, action mistakes, recovery burden, subjective control, and verified time-to-completion differ across materially different users and situations. Where demographic or sensitive attributes are relevant to such research, they should be obtained through voluntary, purpose-specific research consent rather than inferred from names, language, appearance, behavior, or connected data. Crenshaw's critique is especially relevant: a system can appear fair along separate race and gender averages while failing at their intersections (Crenshaw 1989). Buolamwini and Gebru's empirical result shows why subgroup evaluation can materially change what "accuracy" appears to mean (Buolamwini and Gebru 2018). The evaluation design should therefore be capable of detecting compound disadvantage without turning the production system itself into a demographic profiling machine.
The design aspiration is not a digital servant for everyone. That metaphor preserves the very hierarchy the technology could help unsettle. A better aim is portable administrative infrastructure under the user's authorship: enough context to remove repeated explanation, enough evidence discipline to resist confident invention, enough integration to close loops, enough privacy to keep context from becoming exposure, and enough restraint that the user remains the source of ends. Catherine Knight Steele's account of Black feminist technoculture is useful here because it locates Black women within histories of technological skill, creation, entrepreneurship, and communicative expertise rather than treating them only as subjects of technical harm (Steele 2021). The political possibility of a tool is not exhausted by whom it protects from error; it includes whom it equips to act.
That possibility is conditional. FORTH does not abolish administrative burden, unpaid labor, institutional discrimination, inaccessible systems, or unequal access to human help. A personal tool that becomes excellent at navigating a bad bureaucracy can even make the bureaucracy less politically visible. Relief for the individual and reform of the institution are different projects. The defensible ambition is narrower and still significant: when coordination work cannot yet be eliminated at its source, the ability to navigate it competently should not remain scarce simply because competent assistance has historically been costly.
The strongest objection
There is an obvious skeptical reply. Perhaps the “governed layer” is simply another elaborate wrapper around a fast-improving model. If foundation models continue to gain context length, tool use, memory, reasoning, and reliability, why freeze today’s limitations into a product architecture? Why not wait for the model to become the assistant?
The objection is useful because it forces the thesis to separate contingent limitations from enduring requirements. Some FORTH mechanisms may indeed become simpler as models improve. Better speech recognition should reduce dictation ambiguity. Better reasoning should reduce avoidable tool calls. More reliable models may shrink the verification surface for some task classes. Any architecture that treats current weaknesses as eternal would age badly.
The updated system creates a second objection: perhaps the cure for prompt burden is a bureaucracy of Sources. Ten canonical roles, provenance fields, people graphs, voice examples, freshness rules, and evaluation gates can sound like a meticulously organized way of making the user administer the assistant. Worse, a unified context system could become precisely the dossier the personalization argument warns against. This is a serious failure mode, not a rhetorical straw man.
A feminist and Black critical reading adds a third objection that the usual AI-safety vocabulary can miss. The product could take work historically performed by wives, secretaries, domestic workers, care workers, and other disproportionately female or racialized labor forces; rename it "friction"; hide the human histories and contemporary labor still inside the technical stack; and sell the resulting capacity back primarily to affluent professionals. At the same time, the context required to make that system excellent could normalize an unprecedented concentration of intimate personal and professional information. In that version of the future, the assistant would not democratize support. It would automate a service hierarchy, intensify informational asymmetry, and describe both outcomes as personalization. Glenn, Browne, Benjamin, and Steele give different reasons to reject that story, but they converge on the need to analyze labor, classification, access, and power rather than treating technology as an empty vessel (Glenn 1992; Browne 2015; Benjamin 2019; Steele 2021).
Version 0.2’s best answer is architectural restraint. Only Executive Context is a first-use gate, and even it is designed as eleven short, skippable prompts rather than a biography. People, voice calibration, and professional field routing are progressive; a blank or skipped field remains unknown. The Sources are prohibited from becoming copies of live mailboxes, calendars, fares, current news, or account state. Missing specialization is supposed to degrade personalization honestly rather than trigger synthetic completion. Sensitive inferences are barred from silently hardening into durable facts, and writeback is a separate, verifiable act (FORTH Project 2026c, 2026f, 2026g, 2026h, 2026j). These are strong specifications, but they are still specifications. Dogfooding must determine whether the resulting reduction in repeated explanation is worth the setup and maintenance cost.
But several requirements do not disappear with model capability. Current facts still need authoritative sources. The system still needs to know whether an appointment has changed since the last context snapshot. Information flows still require privacy rules. External actions still require authority and auditability. A purchase remains different from a draft because consequences, not tokens, make it different. A user still needs a way to correct durable memory. A system operating in multiple social roles still needs to prevent a fact appropriate in one context from leaking into another. These are protocol and governance problems even when the reasoning engine is excellent.
Nor should integration itself be romanticized. A system with more context can make more contextually appropriate decisions; it can also assemble a more invasive portrait. A system with more action authority can save more labor; it can also cause greater harm. A system that anticipates needs can feel exceptionally competent; it can also become paternalistic. The proper response is not to deny these tensions but to make them testable.
Here the FORTH constitution now goes beyond the germ of falsification. It distinguishes specified, configured, tested, and passing behavior; defines twenty-five adversarial seed cases across dictation, identity, time, privacy, live retrieval, connector failure, action safety, household evidence, coaching, durable correction, and scheduled work; and sets hard release gates that blocked or unrun capabilities are forbidden to satisfy (FORTH Project 2026g). A serious next stage would test FORTH against a strong general-purpose chatbot on repeated, realistic workflows. The outcomes should include time to verified completion; number of user interventions; decisive-fact accuracy; stale-state errors; wrong-recipient and unauthorized-action rates; privacy leakage across personal/executive boundaries; mutation reconciliation after ambiguous provider outcomes; scheduled-task authority escape; correction retention; subjective workload; and measures of whether repeated use increases or diminishes perceived control. Because personalization can increase sycophancy, matched tests with and without user history should also measure factual flips, omitted counterevidence, and excess agreement (Jain et al. 2026; FORTH Project 2026d). Where an ethically designed study can recruit sufficient variation, the same outcomes should be examined across voluntarily self-described gender, race, disability, caregiving, income, age, and other relevant social positions, including their intersections; the production assistant should never infer those research attributes from behavioral traces. The comparison should include adversarial cases in which doing nothing or asking one precise question is the correct behavior.
Such an evaluation could disconfirm the product thesis. Perhaps maintaining context costs more attention than it saves. Perhaps privacy boundaries create so much friction that users prefer separate agents. Perhaps an integrated assistant induces overreliance. Perhaps general models plus ordinary connectors converge on the same performance. Perhaps the tool saves substantial time on average while shifting correction, surveillance, or failure costs onto users with less institutional power. If FORTH cannot outperform a strong baseline on completion while also constraining privacy and action risk—and if its benefits depend on a narrow ideal user—then “chief-of-staff-grade” is a metaphor, not a result.
That possibility is a virtue of the argument, not a weakness. Necessity should not be established by branding. It should be earned by showing that the layer removes a measurable coordination cost that the underlying model does not.
What FORTH is for
The decisive distinction is between generating an answer and maintaining a life in motion.
We already have machines that can produce astonishing intellectual artifacts. Their abundance sharpens Simon’s original diagnosis. The scarce resource is the human capacity to decide what deserves attention, to remember what remains open, to reconcile commitments across contexts, to detect when the world has changed, to separate evidence from inference, and to authorize consequences. Feminist analysis adds that this capacity has never been distributed neutrally: much of the anticipating, remembering, coordinating, caring, and monitoring required to keep households and organizations functional has been made invisible, feminized, or transferred through classed and racialized divisions of labor (Daminger 2019; Glenn 1992). Black feminist analysis adds that there is no honest universal user standing outside race, gender, class, sexuality, disability, family structure, and institutional power (Combahee River Collective 1977; Crenshaw 1989; Collins 1986). Critical race and technology scholarship adds that more data and more automation do not become neutral by becoming computational (Browne 2015; Benjamin 2019).
Those are not reasons to abandon personal AI. They are reasons to specify what kind is worth building. FORTH proposes a bargain in which the person supplies ends, explicit context, corrections, boundaries, and consequential judgment. Versioned Sources hold only the stable context and doctrine worth carrying forward; live systems own the facts that can change; inference remains defeasible; scheduled work carries only the authority deliberately granted to it. The system then absorbs as much surrounding coordination work as can be absorbed safely: repairing intent, resolving people, retrieving current state, filtering options, checking dependencies, planning, drafting, comparing, remembering delegated commitments, preparing artifacts, executing bounded actions, reconciling results, and returning what still requires human judgment. Personal and executive life are coordinated because they share real constraints; they remain distinct because relevance does not imply permission to disclose. Intelligence is coupled to evidence because plausibility is not truth. Action is coupled to authority because capability is not consent. Memory is coupled to correction because personalization is not ownership of the person.
The breadth of the updated system is important because ordinary life is broad. The same person who negotiates a consequential professional decision may later need to research a qualified local clinician, recover a disrupted trip, choose dinner from a photographed refrigerator, verify an electrician, compare two repair estimates, phrase a difficult message, or create a reminder that runs only if an external condition changes. The important technical object is not each vertical in isolation. It is the governed continuity that lets constraints and context travel far enough to make good decisions without letting sensitive information, stale facts, inferred identities, or action authority travel farther than they should (FORTH Project 2026e, 2026g).
This is the strongest argument for treating the category as critical technology. Administrative capacity determines what information becomes action, what right becomes access, what intention survives interruption, what risk is noticed, and what obligation is actually completed. Historically, high levels of that capacity have often been purchased through human labor or supplied invisibly inside families and institutions. Generative AI makes some of the underlying cognition technically reproducible at far lower marginal cost. FORTH's design hypothesis is that governance can make that reproduction operationally trustworthy enough to distribute useful capacity without demanding that the user become a prompt engineer or surrender authorship in exchange for convenience.
The phrase critical technology carries a corresponding obligation. A system that saves executives time while creating a richer apparatus for profiling is not emancipatory. A system that works beautifully for the statistically typical user while failing people at the intersections is not mature. A system that automates the visible click but leaves anticipation, correction, verification, and recovery with the user has not removed the work. A system that helps individuals navigate unjust bureaucracy does not make the bureaucracy just. And a system whose convenience depends on hidden human labor should not represent that labor as machine autonomy. Feminist and Black critical theory do not decorate the product thesis; they supply failure criteria for it.
FORTH is therefore best understood neither as a chatbot nor as a digital servant. It is a proposed governance, context, and coordination layer between generative intelligence and lived systems of action. Its highest ambition is not that the AI should become indispensable. It is that the machinery required to keep a life coherent should consume less conscious attention and less avoidable human labor while leaving the person with more substantive capacity to work, care, decide, relate, recover, and act.
That is authored delegation. Autonomy is not the absence of assistance; it is the preservation of authorship under assistance. Whether FORTH achieves that standard is an empirical question, and the current Sources are specifications rather than proof. But the need they expose is larger than the product: intelligence without context produces content; context without verification produces risk; context without self-definition can become profiling; action without authority produces intrusion; and automation without a theory of power can simply hide where the labor and the harm went.
The missing layer is not more intelligence. It is governed administrative capacity. FORTH deserves to become critical technology only if it can make that capacity more available while reducing—not redistributing—burden, surveillance, and loss of control.
Research note: source discipline
This manuscript uses sources according to the job each can legitimately perform.
Interdisciplinary reference council. The FORTH Collegium on Human Agency & Intelligent Systems is an organizing device for adversarial reading across cognition, HCI, automation, sociology of work, feminist and Black feminist theory, critical race and technology studies, privacy, political philosophy, and human–AI performance. Its principal intellectual reference points in this version include Herbert Simon, Lucy Suchman, Edwin Hutchins, Lisanne Bainbridge, Allison Daminger, Evelyn Nakano Glenn, Kimberlé Crenshaw, Patricia Hill Collins, Joan Tronto, Helen Nissenbaum, Simone Browne, Ruha Benjamin, Catherine D'Ignazio, Lauren Klein, Catherine Knight Steele, Amartya Sen, and the empirical human–AI researchers cited below. None of these scholars participated in, reviewed, advised, or endorsed FORTH. "Council" names a method of putting their arguments into disciplined tension, not an actual advisory relationship.
Product primary documents. The FORTH constitution, onboarding/context templates, and operating playbooks are primary documentary evidence of the system's intended design, doctrine, boundaries, capability envelope, and evaluation logic. They are not evidence that the design has been implemented successfully or that it produces claimed outcomes. A blank onboarding template is evidence of the intended schema, not evidence about any user; a playbook is not evidence that its connector is currently available. Accordingly, the essay uses formulations such as "specifies," "proposes," and "adopts" rather than treating design intent as performance evidence.
Primary empirical research. Beck and Jackson (2022), Noy and Zhang (2023), Brynjolfsson, Li, and Raymond (2025), Buçinca, Malaya, and Gajos (2021), Buolamwini and Gebru (2018), Dell'Acqua et al. (2026), Daminger (2019), Greshake et al. (2023), Grinschgl, Papenmeier, and Meyerhoff (2021), Jain et al. (2026), Rocha Beardall, Mueller, and Cheng (2024), and Rubinstein, Meyer, and Evans (2001) report original empirical research or demonstrations. Quantitative effects are confined to the populations, tasks, systems, and outcomes those studies actually examined. Qualitative findings, security demonstrations, and subgroup-disparity results are used only for the narrower propositions their designs can support; none validates FORTH.
Secondary synthesis. Risko and Gilbert (2016), Gilbert et al. (2023), Onnasch et al. (2014), Vaccaro, Almaatouq, and Malone (2024), and Passmore, Olafsson, and Tee (2026) synthesize literatures. Their role is to establish the state or pattern of a research field, not to provide product-specific validation.
Feminist, Black feminist, and critical race/technology sources. The Combahee River Collective Statement (1977) is a primary historical political text. Crenshaw (1989) and Collins (1986) are foundational legal and sociological works used for intersectional analysis and self-definition, not empirical product claims. Glenn (1992) supplies historical sociology of racialized and gendered reproductive labor; Tronto (1993) supplies a political ethic of care. Browne (2015), Benjamin (2019), D'Ignazio and Klein (2020), Ray, Herd, and Moynihan (2023), and Steele (2021) provide historical, theoretical, or critical analyses of surveillance, race, data, administrative burden, labor, power, and technology. Their role is diagnostic and normative: they identify questions a consequential system must answer. The manuscript does not imply that any of these authors anticipated, endorsed, or empirically validated FORTH.
Other foundational and official sources. Simon (1971), Suchman (1987), Hutchins (1995), Nissenbaum (2004, 2019), Illich (1973), Ryan and Deci (2006), and Sen (1999) provide conceptual or normative frameworks. The essay does not convert those frameworks into causal efficacy claims. NIST (2024) is official risk-management guidance and is cited as such, not as empirical proof.
This separation matters because citation quality is not a matter of quantity. A citation is correct only when the evidentiary type, population, historical scope, and inference match the claim it is asked to support. Critical theory is not a substitute for efficacy evidence; efficacy evidence is not a substitute for a theory of power.
References
Bainbridge, Lisanne. 1983. “Ironies of Automation.” Automatica 19 (6): 775–779. https://doi.org/10.1016/0005-1098(83)90046-8.
Beck, Emorie D., and Joshua J. Jackson. 2022. “Personalized Prediction of Behaviors and Experiences: An Idiographic Person–Situation Test.” Psychological Science 33 (10): 1767–1782. https://doi.org/10.1177/09567976221093307.
Benjamin, Ruha. 2019. Race After Technology: Abolitionist Tools for the New Jim Code. Cambridge, UK: Polity. https://www.ruhabenjamin.com/race-after-technology.
Browne, Simone. 2015. Dark Matters: On the Surveillance of Blackness. Durham, NC: Duke University Press. https://doi.org/10.1215/9780822375302.
Brynjolfsson, Erik, Danielle Li, and Lindsey Raymond. 2025. “Generative AI at Work.” The Quarterly Journal of Economics 140 (2): 889–942. https://doi.org/10.1093/qje/qjae044.
Buçinca, Zana, Maja Barbara Malaya, and Krzysztof Z. Gajos. 2021. “To Trust or to Think: Cognitive Forcing Functions Can Reduce Overreliance on AI in AI-Assisted Decision-Making.” Proceedings of the ACM on Human-Computer Interaction 5 (CSCW1), Article 188. https://doi.org/10.1145/3449287.
Buolamwini, Joy, and Timnit Gebru. 2018. “Gender Shades: Intersectional Accuracy Disparities in Commercial Gender Classification.” In Proceedings of the 1st Conference on Fairness, Accountability and Transparency, 77–91. Proceedings of Machine Learning Research 81. https://proceedings.mlr.press/v81/buolamwini18a.html.
Collins, Patricia Hill. 1986. “Learning from the Outsider Within: The Sociological Significance of Black Feminist Thought.” Social Problems 33 (6): S14–S32. https://doi.org/10.2307/800672.
Combahee River Collective. 1977. “The Combahee River Collective Statement.” April. Library of Congress Web Archive. https://www.loc.gov/item/lcwaN0028151/.
Crenshaw, Kimberlé. 1989. “Demarginalizing the Intersection of Race and Sex: A Black Feminist Critique of Antidiscrimination Doctrine, Feminist Theory and Antiracist Politics.” University of Chicago Legal Forum 1989, Article 8. https://chicagounbound.uchicago.edu/uclf/vol1989/iss1/8/.
Daminger, Allison. 2019. “The Cognitive Dimension of Household Labor.” American Sociological Review 84 (4): 609–633. https://doi.org/10.1177/0003122419859007.
Dell’Acqua, Fabrizio, Edward McFowland III, Ethan Mollick, Hila Lifshitz, Katherine C. Kellogg, Saran Rajendran, Lisa Krayer, François Candelon, and Karim R. Lakhani. 2026. “Navigating the Jagged Technological Frontier: Field Experimental Evidence of the Effects of Artificial Intelligence on Knowledge Worker Productivity and Quality.” Organization Science 37 (2): 403–423. https://doi.org/10.1287/orsc.2025.21838.
D’Ignazio, Catherine, and Lauren F. Klein. 2020. Data Feminism. Cambridge, MA: MIT Press. https://data-feminism.mitpress.mit.edu/.
FORTH Project. 2026a. Dictation Intent: Progressive Broken-STT Correction Form and User-Specific Intent Reconstruction Model. Project Source 04, schema version 0.2, August 7. Unpublished onboarding and dictation-reconstruction specification supplied by the author.
FORTH Project. 2026b. Email + Calendar. Project Source 05, version 0.2, August 7. Unpublished operating playbook supplied by the author.
FORTH Project. 2026c. Executive Context: First-Use Onboarding Form and Canonical User-Context Template. Project Source 01, schema version 0.2, August 7. Unpublished onboarding specification supplied by the author.
FORTH Project. 2026d. Human Performance Coaching — Executive Psychology, Neuroscience, Philosophy & Adaptive Human Performance. Project Source 09, source version 0.5, August 7. Unpublished operating playbook supplied by the author.
FORTH Project. 2026e. Life Logistics — Food, Household Operations & Local Services Playbook. Project Source 08, version 2.0, August 7. Unpublished operating playbook supplied by the author.
FORTH Project. 2026f. People Relationships: Progressive Onboarding Form and Canonical People/Group Identity Template. Project Source 02, schema version 0.2, August 7. Unpublished onboarding specification supplied by the author.
FORTH Project. 2026g. Personal + Executive Assistant — MVP Behavioral DNA and Source Architecture. Product Constitution, version 0.2, August 7. Unpublished design specification supplied by the author.
FORTH Project. 2026h. Research Field Pack: Progressive Professional Research Bootstrap and Canonical Field-Specific Evidence-Routing Template. Project Source 07, schema version 0.2, August 7. Unpublished onboarding and research-routing specification supplied by the author.
FORTH Project. 2026i. Travel & Logistics Playbook. Project Source 06, version 1.1, August 7. Unpublished operating playbook supplied by the author.
FORTH Project. 2026j. Voice Correspondence: Progressive Voice Calibration Form and Canonical Correspondence-Style Baseline. Project Source 03, schema version 0.2, August 7. Unpublished onboarding and correspondence-style specification supplied by the author.
Gilbert, Sam J., Annika Boldt, Chhavi Sachdeva, Chiara Scarampi, and Pei-Chun Tsai. 2023. “Outsourcing Memory to External Tools: A Review of ‘Intention Offloading.’” Psychonomic Bulletin & Review 30 (1): 60–76. https://doi.org/10.3758/s13423-022-02139-4.
Glenn, Evelyn Nakano. 1992. “From Servitude to Service Work: Historical Continuities in the Racial Division of Paid Reproductive Labor.” Signs: Journal of Women in Culture and Society 18 (1): 1–43. https://doi.org/10.1086/494777.
Greshake, Kai, Sahar Abdelnabi, Shailesh Mishra, Christoph Endres, Thorsten Holz, and Mario Fritz. 2023. “Not What You’ve Signed Up For: Compromising Real-World LLM-Integrated Applications with Indirect Prompt Injection.” In Proceedings of the 2023 ACM Workshop on Artificial Intelligence and Security, 79–90. New York: ACM. https://doi.org/10.1145/3605764.3623985.
Grinschgl, Sandra, Frank Papenmeier, and Hauke S. Meyerhoff. 2021. “Consequences of Cognitive Offloading: Boosting Performance but Diminishing Memory.” Quarterly Journal of Experimental Psychology 74 (9): 1477–1496. https://doi.org/10.1177/17470218211008060.
Hutchins, Edwin. 1995. Cognition in the Wild. Cambridge, MA: MIT Press. https://mitpress.mit.edu/9780262581462/cognition-in-the-wild/.
Illich, Ivan. 1973. Tools for Conviviality. New York: Harper & Row.
Jain, Shomik, Charlotte Park, Matt Viana, Ashia Wilson, and Dana Calacci. 2026. “Interaction Context Often Increases Sycophancy in LLMs.” In Proceedings of the 2026 CHI Conference on Human Factors in Computing Systems, Article 793, 1–26. New York: ACM. https://doi.org/10.1145/3772318.3791915.
National Institute of Standards and Technology (NIST). 2024. Artificial Intelligence Risk Management Framework: Generative Artificial Intelligence Profile. NIST AI 600-1. Gaithersburg, MD: National Institute of Standards and Technology. https://doi.org/10.6028/NIST.AI.600-1.
Nissenbaum, Helen. 2004. “Privacy as Contextual Integrity.” Washington Law Review 79 (1): 119–158. https://digitalcommons.law.uw.edu/wlr/vol79/iss1/10/.
Nissenbaum, Helen. 2019. “Contextual Integrity Up and Down the Data Food Chain.” Theoretical Inquiries in Law 20 (1): 221–256. https://doi.org/10.1515/til-2019-0008.
Noy, Shakked, and Whitney Zhang. 2023. “Experimental Evidence on the Productivity Effects of Generative Artificial Intelligence.” Science 381 (6654): 187–192. https://doi.org/10.1126/science.adh2586.
Onnasch, Linda, Christopher D. Wickens, Huiyang Li, and Dietrich Manzey. 2014. “Human Performance Consequences of Stages and Levels of Automation: An Integrated Meta-Analysis.” Human Factors 56 (3): 476–488. https://doi.org/10.1177/0018720813501549.
Passmore, Jonathan, Bergsveinn Olafsson, and David Tee. 2026. “A Systematic Literature Review of Artificial Intelligence (AI) in Coaching: Insights for Future Research and Product Development.” Journal of Work-Applied Management 18 (1): 110–129. https://doi.org/10.1108/JWAM-11-2024-0164.
Ray, Victor, Pamela Herd, and Donald Moynihan. 2023. “Racialized Burdens: Applying Racialized Organization Theory to the Administrative State.” Journal of Public Administration Research and Theory 33 (1): 139–152. https://doi.org/10.1093/jopart/muac001.
Risko, Evan F., and Sam J. Gilbert. 2016. “Cognitive Offloading.” Trends in Cognitive Sciences 20 (9): 676–688. https://doi.org/10.1016/j.tics.2016.07.002.
Rocha Beardall, Theresa, Collin Mueller, and Tony Cheng. 2024. “Intersectional Burdens: How Social Location Shapes Interactions with the Administrative State.” RSF: The Russell Sage Foundation Journal of the Social Sciences 10 (4): 84–102. https://doi.org/10.7758/RSF.2024.10.4.04.
Rubinstein, Joshua S., David E. Meyer, and Jeffrey E. Evans. 2001. “Executive Control of Cognitive Processes in Task Switching.” Journal of Experimental Psychology: Human Perception and Performance 27 (4): 763–797. https://doi.org/10.1037/0096-1523.27.4.763.
Ryan, Richard M., and Edward L. Deci. 2006. “Self-Regulation and the Problem of Human Autonomy: Does Psychology Need Choice, Self-Determination, and Will?” Journal of Personality 74 (6): 1557–1585. https://doi.org/10.1111/j.1467-6494.2006.00420.x.
Sen, Amartya. 1999. Development as Freedom. New York: Alfred A. Knopf.
Simon, Herbert A. 1971. “Designing Organizations for an Information-Rich World.” In Computers, Communications, and the Public Interest, edited by Martin Greenberger, 37–72. Baltimore: Johns Hopkins Press.
Steele, Catherine Knight. 2021. Digital Black Feminism. New York: NYU Press. https://nyupress.org/9781479808380/digital-black-feminism/.
Suchman, Lucy A. 1987. Plans and Situated Actions: The Problem of Human-Machine Communication. Cambridge: Cambridge University Press.
Tronto, Joan C. 1993. Moral Boundaries: A Political Argument for an Ethic of Care. New York: Routledge.
Vaccaro, Michelle, Abdullah Almaatouq, and Thomas W. Malone. 2024. “When Combinations of Humans and AI Are Useful: A Systematic Review and Meta-Analysis.” Nature Human Behaviour 8: 2293–2303. https://doi.org/10.1038/s41562-024-02024-1.
FORTH
Explore the system examined in this essay
FORTH is the governed Personal + Executive intelligence layer discussed here: a design for coordinating durable human context, live external truth, bounded action authority, and cross-domain follow-through while preserving privacy, correction, and user authorship.