Skip to content

chore(operations): prove external hourly scheduler continuation and error recovery #96

Description

@seonghobae

Problem

Noema's repository-owned maintenance/product-development controls cannot prove that the external ChatGPT hourly task is enabled, contains the current compact execution prompt, survives a generic scheduled-task failure, or continues past one useful action. A scheduler prompt edit, chat claim, or generic There was a problem with your scheduled task response is not repository execution evidence.

This issue owns the external control-plane evidence gap identified as G-17 by the canonical documentation audit on PR #71. It does not duplicate repository workflow implementation owned by #80 or live App/ruleset provisioning owned by #27/#29.

Protected main observed when opening this issue: c85d710804139c0697d7ef8fa47d02b1389e6d84.

Compact external task prompt

The external scheduler should keep one enabled hourly task and use a compact prompt that delegates stable detail to repository authority:

Continuously improve ContextualWisdomLab/noema toward defensible commercial/acquisition readiness. Execute, do not merely report. Start from fresh protected-main, every open PR/issue, exact heads/live bases, stacks, reviews/threads/checks/security/rules/releases/docs and active-writer evidence; rebuild after every material action. Treat pending, absent, skipped, stale, predecessor, synthetic, model-only, rate-limited or status-only evidence as non-passing.

Write only Noema. Before every write refetch exact target/base/blob/ref/review/writer state; freeze only raced branches and rotate. Never force-push, self-approve, weaken gates, fabricate authority/secrets/evidence, or create repair/self-modifying branch-patching workflows.

Priority: merge only unchanged gate-clean authorized PRs; test-first fix current product/security/reliability/data/accessibility defects; remove Noema-owned blockers; resolve only addressed threads/duplicates; advance stacks/issues; run protected-main operational acceptance; repair canonical docs and executable contracts; convert gaps into source/tests/operators; then implement the highest-impact bounded buyer slice. After every action or defer, return to queue top.

Use RCA -> distinct remedies -> feasibility -> smallest safe action -> exact proof. Waiting blocks only that exact lane. Prompt repair, inventory, docs, one test, one commit, one PR update, one merge or one blocker is intermediate. After user redirection, perform at least two materially distinct repository actions when two safe lanes exist. Never end on test-only RED while safe GREEN exists. Documentation must hand off to the highest-priority safe non-documentation action.

On a generic scheduled-task error, refetch this task and GitHub, keep one enabled hourly task, simplify this prompt if needed, do not invent hidden error codes, and immediately resume repository execution. Stable architecture/security/product detail belongs in AGENTS.md and canonical PRD/TRD/ARCHITECTURE/ADR/UML/ERD/TRACEABILITY/security/test/operability/licensing authority.

Before exit, perform two consecutive fresh whole-Noema sweeps. Any safe merge, mutation, test, closure, stack repair, operational proof, docs repair, release preparation or bounded product action resets the sweep count. End only on practical invocation-budget exhaustion or two clean sweeps proving every remaining lane non-actionable. Routine status remains internal.

Acceptance criteria

Scheduler identity and configuration

  • Record one enabled hourly task identity, owner, timezone/schedule, observation time and prompt digest in access-controlled operational evidence.
  • Prove duplicate/obsolete Noema scheduler tasks are disabled rather than concurrently writing.
  • Bind the task to ContextualWisdomLab/noema only; .github, naruon, contextual-orchestrator and other dedicated-writer repositories remain read-only dependencies.
  • Confirm stable product/architecture/security detail is referenced from repository authority rather than copied into diverging scheduler prose.

Generic-error recovery

  • Exercise or observe one bounded generic task failure without inventing a hidden provider error code.
  • Prove the next invocation refetches task state and GitHub live state before acting.
  • Prove prompt repair earns zero completion credit and repository execution resumes in the same invocation when a safe lane exists.
  • Retain failure timestamp, scheduler/task identity, next successful invocation identity and exact resulting GitHub mutation(s).

Work-conserving behavior

  • On a run with two safe independent lanes, retain evidence of at least two materially distinct repository actions.
  • Prove a waiting CI/review lane rotates to another safe lane rather than ending the run.
  • Prove documentation work hands off to a non-documentation action when one is safe.
  • Prove test-first work does not stop on RED when a safe GREEN implementation exists.
  • Retain two fresh exit sweeps or an explicit practical invocation-budget boundary.

Writer safety and authority

  • Prove exact pre-write head/base/target/review/writer refetch and branch-local freeze on detected competing writes.
  • Prove no force push, self-approval, gate weakening, synthetic evidence, repair workflow or credential fallback occurs.
  • Keep external task evidence separate from GitHub checks, formal reviews, protected merge authority, release/deployment acceptance and acquisition evidence.

Evidence schema (minimum)

Retain a bounded record with descriptive multiword snake_case fields:

{
  "schema_version": 1,
  "scheduler_task_identity": "provider-scoped opaque id",
  "prompt_sha256": "64 lowercase hex",
  "scheduled_at": "ISO-8601 UTC",
  "started_at": "ISO-8601 UTC",
  "repository_full_name": "ContextualWisdomLab/noema",
  "protected_main_sha": "40 lowercase hex",
  "generic_error_observed": false,
  "github_actions_performed": [],
  "deferred_lanes": [],
  "exit_sweep_count": 2,
  "remaining_non_actionable_reasons": []
}

The record must not contain secrets, private keys, raw tokens, hidden model reasoning, vulnerability details or unnecessary personal data.

Documentation handoff

PR #71 remains the canonical PRD/TRD/Architecture/ADR/UML/ERD/Traceability owner. Reconcile this issue into canonical traceability/operability after #71 stabilizes; do not create a parallel architecture authority. PR #80 remains the source implementation owner for repository-side work-conserving scheduler/publisher behavior.

Non-goals

Related: #27, #29, #71, #80

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    area: authAuthentication, authorization, identity, or tenant isolationarea: ci-cdCI, GitHub Actions, checks, release, or supply chainarea: dependenciesDependency or lockfile maintenancearea: securitySecurity boundary, hardening, or vulnerability preventionpriority: mediumNormal-priority or P2 workscope: researchResearch, statistical validation, or scientific evidencestatus: triagedOpen issue has an organization taxonomy assignmenttype: maintenanceMaintenance, build, dependency, or operational upkeep

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions