The alignment problem, and how CORe solves it
An engineering overview of COReAlign, the stack of architectural constraints, prompt pushing, and rulesets that keeps CORe's models honest, transparent, and structurally safe.
1. The alignment problem #
Alignment is the problem of getting a capable model to do what its users and operators actually want, not what a clever reward signal, filter, or prompt appears to ask for. The failure modes aren't exotic: a well-trained model will confidently hallucinate, silently follow a jailbreak, or optimize for the shape of a helpful answer instead of the substance.
Most production mitigations bolt alignment on after the model is trained. A toxicity classifier screens outputs. A system prompt tells the model to refuse. A moderation endpoint blocks the tokens that got past both. These controls are policy, not architecture. Under adversarial pressure, policy bends.
2. Why traditional safety filters fail #
Post-hoc filters share three structural weaknesses:
- They run at the output boundary. The model has already reasoned its way to the unsafe answer; the filter only decides whether to ship it. That means the unsafe reasoning exists internally and can leak through paraphrase, tool-use, or partial-output streaming.
- They classify tokens, not intent. A classifier doesn't know whether the user is a security researcher demonstrating an exploit or an attacker writing one. It pattern-matches surface features.
- They drift as prompts drift. Every jailbreak template ever published is a demonstration that a policy layer ungrounded from the base model is probing-distance away from being bypassed.
3. COReAlign: alignment at the architecture level #
COReAlign is the layered alignment system that wraps every CORe model. It is designed around a simple premise: the earlier in the stack alignment is enforced, the harder it is to bypass from outside. We build alignment into the forward pass, not onto it.
3.1 Training-time objectives
The base model is trained against a joint objective: next-token prediction plus an alignment loss derived from a curated set of counterfactual preferences. Rather than "prefer A to B," we train "prefer A, and explain why B would have been wrong", so the model learns to represent the reason, not just the choice.
3.2 Architectural constraints
A small number of attention heads in each block are designated suppressors. They are gated so that when a suppressor activates above a threshold, a corresponding set of output distributions is renormalized to zero mass on a narrow token class. These gates are learned, not hand-coded, but because they're part of the architecture, they travel with the weights.
4. Prompt pushing #
Traditional system prompts are just more user tokens. If they sit in the same stream that a user controls, a user who is clever enough can negotiate with them. Prompt pushing sidesteps this by injecting directive vectors directly into the residual stream at designated layers, rather than into the token sequence.
Operationally, each directive is compiled into a sparse delta on the residual stream, a tensor the model was trained to treat as context-with-authority. From the user's perspective, nothing in the chat transcript reveals the directive; from the model's perspective, it cannot be overwritten by later user tokens because it never competed for attention weight in the first place.
# pseudo-code: attach a pushed directive to a generation call
from corealign import Directive, push
directive = Directive(
id="no-exfil-credentials",
scope="operator",
priority=10,
layers=[14, 22, 31],
)
response = model.generate(
prompt=user_message,
pushes=[push(directive)],
)
Directives carry a scope (operator, platform,
user), a priority, and a layer-set. They are applied in
priority order, and lower-scope directives cannot override
higher-scope ones; a user can't push the platform around.
5. Rulesets #
Rulesets are the declarative side of the stack. Where prompt pushing shapes the model's internal context, rulesets define external contracts that every generation is checked against: live, before the output leaves the model runner.
A ruleset is authored in a compact DSL, compiled to a byte-coded check, and versioned like code. Every production deploy pins an exact ruleset hash, so alignment behavior is reproducible across model versions.
# ruleset: core-6/public/v7
rule "honest-uncertainty":
when response.confidence < 0.6
must response.contains_hedge == true
else revise("add explicit uncertainty")
rule "no-fabricated-citations":
when response.has_citations
must every(c in response.citations: c.resolvable)
else strip_unresolvable_citations()
rule "operator-scope":
when directive.scope == "operator"
must response.respects(directive)
else refuse(reason=directive.reason)
Rulesets are where policy lives, and that's intentional: policy changes faster than model weights, and we don't want to retrain every time a product decision shifts. Rules compile down to fast checks (single-digit milliseconds per generation) and can be audited end-to-end.
5.1 Tiers
| Tier | Editable by | Can override |
|---|---|---|
| Platform | CORe engineering | (root tier) |
| Operator | Enterprise admin | User rules |
| User | End user / prompt | None |
When a lower-tier rule tries to relax a higher-tier constraint, the generation is refused with the higher tier's reason surfaced to the caller. No silent softening.
6. Transparent reasoning #
CORe models emit reasoning in a dedicated channel separate from their final answer. Two properties fall out of this:
- Refusals are explainable. When a ruleset triggers, the reasoning channel captures the specific rule and cited evidence.
- Hallucinations are catchable. If the reasoning channel cites sources, each citation is checked by a rule; fabricated references are stripped before the answer ships.
7. Evaluation & results #
COReAlign is evaluated on a combination of internal suites and
public benchmarks. The headline numbers for CORe 6.1:
| Benchmark | Metric | Score |
|---|---|---|
| HarmBench (adversarial) | Refusal-appropriate rate | 99.0% |
| TruthfulQA | Truthful (mc1) | 88.8% |
| MMLU | Capability (kept high) | 91.8% |
The point of that last row is worth stating plainly: alignment does not have to cost capability. Models that hallucinate less and refuse more precisely also reason more cleanly on the benchmarks their users actually care about.
8. What's next #
The next work on COReAlign is twofold: extending prompt pushing to multimodal inputs, and opening up the ruleset compiler so operators can author and audit their own rules without a CORe engineer in the loop.
If you want to go deeper, the best place to start is the chat app; every response is produced under the stack described above, and the reasoning channel is visible to you.