CORe Docs
Home Chat Try CORe
CORe / Docs / Alignment

The alignment problem, and how CORe solves it

An engineering overview of COReAlign, the stack of architectural constraints, prompt pushing, and rulesets that keeps CORe's models honest, transparent, and structurally safe.

Topic · Alignment Applies to · CORe 5, CORe 6/6.1 Updated · April 2026

1. The alignment problem #

Alignment is the problem of getting a capable model to do what its users and operators actually want, not what a clever reward signal, filter, or prompt appears to ask for. The failure modes aren't exotic: a well-trained model will confidently hallucinate, silently follow a jailbreak, or optimize for the shape of a helpful answer instead of the substance.

Most production mitigations bolt alignment on after the model is trained. A toxicity classifier screens outputs. A system prompt tells the model to refuse. A moderation endpoint blocks the tokens that got past both. These controls are policy, not architecture. Under adversarial pressure, policy bends.

The short version Filters tell a model what not to say. COReAlign shapes what it can compute in the first place.

2. Why traditional safety filters fail #

Post-hoc filters share three structural weaknesses:

3. COReAlign: alignment at the architecture level #

COReAlign is the layered alignment system that wraps every CORe model. It is designed around a simple premise: the earlier in the stack alignment is enforced, the harder it is to bypass from outside. We build alignment into the forward pass, not onto it.

1
Training-time objectives
Constitutional losses and counterfactual preference data baked into the base model.
2
Architectural constraints
Routing gates and head-level suppressors that structurally limit a class of harmful outputs.
3
Prompt pushing
Immutable directives injected into the residual stream rather than the input tokens.
4
Rulesets
Declarative, versioned policies compiled into fast runtime checks around every generation.
5
Transparent reasoning
Model emits its internal reasoning separately, so refusals and answers are both auditable.

3.1 Training-time objectives

The base model is trained against a joint objective: next-token prediction plus an alignment loss derived from a curated set of counterfactual preferences. Rather than "prefer A to B," we train "prefer A, and explain why B would have been wrong", so the model learns to represent the reason, not just the choice.

3.2 Architectural constraints

A small number of attention heads in each block are designated suppressors. They are gated so that when a suppressor activates above a threshold, a corresponding set of output distributions is renormalized to zero mass on a narrow token class. These gates are learned, not hand-coded, but because they're part of the architecture, they travel with the weights.

Why this matters A filter can be unplugged. A suppressor head is part of the model's forward pass; removing it degrades the model's capability, which means the safety property isn't free to strip.

4. Prompt pushing #

Traditional system prompts are just more user tokens. If they sit in the same stream that a user controls, a user who is clever enough can negotiate with them. Prompt pushing sidesteps this by injecting directive vectors directly into the residual stream at designated layers, rather than into the token sequence.

Operationally, each directive is compiled into a sparse delta on the residual stream, a tensor the model was trained to treat as context-with-authority. From the user's perspective, nothing in the chat transcript reveals the directive; from the model's perspective, it cannot be overwritten by later user tokens because it never competed for attention weight in the first place.

# pseudo-code: attach a pushed directive to a generation call
from corealign import Directive, push

directive = Directive(
    id="no-exfil-credentials",
    scope="operator",
    priority=10,
    layers=[14, 22, 31],
)

response = model.generate(
    prompt=user_message,
    pushes=[push(directive)],
)

Directives carry a scope (operator, platform, user), a priority, and a layer-set. They are applied in priority order, and lower-scope directives cannot override higher-scope ones; a user can't push the platform around.

5. Rulesets #

Rulesets are the declarative side of the stack. Where prompt pushing shapes the model's internal context, rulesets define external contracts that every generation is checked against: live, before the output leaves the model runner.

A ruleset is authored in a compact DSL, compiled to a byte-coded check, and versioned like code. Every production deploy pins an exact ruleset hash, so alignment behavior is reproducible across model versions.

# ruleset: core-6/public/v7
rule "honest-uncertainty":
  when  response.confidence < 0.6
  must  response.contains_hedge == true
  else  revise("add explicit uncertainty")

rule "no-fabricated-citations":
  when  response.has_citations
  must  every(c in response.citations: c.resolvable)
  else  strip_unresolvable_citations()

rule "operator-scope":
  when  directive.scope == "operator"
  must  response.respects(directive)
  else  refuse(reason=directive.reason)

Rulesets are where policy lives, and that's intentional: policy changes faster than model weights, and we don't want to retrain every time a product decision shifts. Rules compile down to fast checks (single-digit milliseconds per generation) and can be audited end-to-end.

5.1 Tiers

TierEditable byCan override
PlatformCORe engineering(root tier)
OperatorEnterprise adminUser rules
UserEnd user / promptNone

When a lower-tier rule tries to relax a higher-tier constraint, the generation is refused with the higher tier's reason surfaced to the caller. No silent softening.

6. Transparent reasoning #

CORe models emit reasoning in a dedicated channel separate from their final answer. Two properties fall out of this:

  1. Refusals are explainable. When a ruleset triggers, the reasoning channel captures the specific rule and cited evidence.
  2. Hallucinations are catchable. If the reasoning channel cites sources, each citation is checked by a rule; fabricated references are stripped before the answer ships.

7. Evaluation & results #

COReAlign is evaluated on a combination of internal suites and public benchmarks. The headline numbers for CORe 6.1:

BenchmarkMetricScore
HarmBench (adversarial)Refusal-appropriate rate99.0%
TruthfulQATruthful (mc1)88.8%
MMLUCapability (kept high)91.8%

The point of that last row is worth stating plainly: alignment does not have to cost capability. Models that hallucinate less and refuse more precisely also reason more cleanly on the benchmarks their users actually care about.

8. What's next #

The next work on COReAlign is twofold: extending prompt pushing to multimodal inputs, and opening up the ruleset compiler so operators can author and audit their own rules without a CORe engineer in the loop.

If you want to go deeper, the best place to start is the chat app; every response is produced under the stack described above, and the reasoning channel is visible to you.