Skip to main content

· 4 min read

The cluster ran on kubectl apply, one-off Helm commands and a handful of scripts for long enough that nobody could say what the actual state was. The manifests existed somewhere, secrets lived in shell history and password managers, and rolling something back meant remembering what it looked like before. This is how the cluster moved to a Git repository as its source of truth.

The goal​

After the migration: the cluster is the output of a repository. A change is a commit, reviewable, revertible, and reconciled automatically. Manual kubectl apply is the exception, not the workflow.

Tooling​

  • Flux CD as the GitOps controller, bootstrapped into the cluster.
  • SOPS with age for encrypting secrets in the repository.
  • flux check --pre before bootstrapping, to validate the cluster meets the prerequisites.

The bootstrap installs the Flux controllers and wires them to the repository, after which everything else arrives through reconciliation.

Repository layout​

clusters/production/     # flux-system + one Kustomization per app
apps/<app>/base/ # manifests for a single app
apps/<app>/kustomization.yaml
apps/<app>/secrets.enc.yaml
infrastructure/namespaces/
scripts/
docs/

The clusters/ directory describes what should exist in the cluster; apps/ describes what each thing is; infrastructure/ holds shared prerequisites like namespaces. Splitting them matters because an app's manifests should be movable without changing how the cluster consumes them.

How reconciliation is wired​

Each app gets its own Flux Kustomization resource that points at its directory:

  • interval: 10m — how often the repository is compared to the cluster.
  • retryInterval: 2m, timeout: 5m — bounded retries for a broken apply.
  • prune: true — resources removed from Git are removed from the cluster.
  • wait: true plus health checks — reconciliation is not "done" until the resources are actually healthy.
  • decryption.provider: sops — secrets are decrypted in-cluster at apply time.

Ten minutes sounds slow when iterating, and it is. For a home cluster the trade is fine: drift gets corrected without anyone watching, and a bad commit is undone with git revert instead of a manual cleanup.

Secrets​

Secrets are committed encrypted. SOPS is configured with an age key pair: the public key is used to encrypt, the private key never enters Git. The private key is backed up offline, and the cluster receives it once as a sops-age secret in the flux-system namespace so the controllers can decrypt at apply time.

The practical benefit is reviewability: an encrypted diff still shows which keys changed, so a secret rotation is visible in a pull request without ever exposing the value.

Proving it with a pilot app​

The migration was validated with a single representative app — a web application with a database, an encrypted secret, persistent storage and an ingress. Once that app reconciled end to end, the pattern was repeated for the rest.

Verification steps that were worth formalizing:

  • flux get kustomizations shows Ready and Applied revision.
  • kubectl get all -n <app> matches the repository.
  • An intentional annotation change in Git shows up in the cluster.
  • git revert of that change rolls it back without manual intervention.

What stayed manual​

Node-level configuration (the kubelet and datastore settings on each control-plane node) is not part of this repository. GitOps covers workloads and their configuration, not the machines running them. That boundary is worth writing down so it is not mistaken for drift.

Takeaways​

  • The win is not automation for its own sake; it is that the intended state is written down once and reviewed like code.
  • Encrypted secrets in Git are workable when the key management is explicit and the private key is backed up somewhere that is not the repository.
  • A pilot app is enough to prove the layout before migrating everything.
  • Keep a rollback path (git revert) and test it early, while the change is still small.

· 5 min read

The symptom​

A host on the home LAN became unreachable from other LAN devices. Every ping and connection attempt failed, while the same host answered normally over its Tailscale address. The router logged ICMP redirects toward the affected host during the failures, which pointed at a routing disagreement rather than a broken interface.

How Tailscale subnet routers work​

A subnet router advertises routes for a physical network into the tailnet, so remote clients can reach that network without running Tailscale on every device. Nodes that accept those routes install them in a separate routing table and add a policy rule that sends matching traffic through the Tailscale interface instead of the default route.

That design is what makes subnet routing convenient — and what makes overlaps dangerous.

Root cause​

One host on the LAN was running Tailscale as a subnet router for the same subnet it was connected to, with route acceptance enabled. Two facts combined:

  • Inbound traffic reached the host over the LAN, as expected.
  • Replies to that traffic matched the accepted-route rule and left through the Tailscale interface.

The return path no longer matched the request path. Requests arrived over Ethernet, replies departed over the VPN, and the peer discarded the replies because they never came back the way they went out. This is asymmetric routing: every interface is up, every route looks plausible in isolation, and traffic still disappears.

The detail that makes this nasty is that the host is the only one affected. Other LAN devices route to it normally; the problem lives entirely in its policy routing rules.

A useful contrast is how different systems ship: some appliance operating systems install a high-priority rule that keeps local-network destinations in the main routing table by default, while a plain general-purpose Linux install does not. The same network, the same Tailscale settings, different failure behavior — the protection rule is the difference.

The fix​

Keep traffic destined for the local subnet in the main routing table, with a priority higher than the accepted-routes rule:

ip rule add from all to <lan-subnet> table main priority 5000

To survive reboots, persist it as a small systemd unit:

[Unit]
Description=Keep local subnet traffic in the main routing table
After=network-online.target tailscaled.service

[Service]
Type=oneshot
ExecStart=/usr/sbin/ip rule add from all to <lan-subnet> table main priority 5000
RemainAfterExit=yes

[Install]
WantedBy=multi-user.target

Verify by checking which table wins for a lookup toward another LAN host:

ip rule show
ip route get <lan-host> from <affected-host>

Before the fix, the lookup resolves through the Tailscale table; after it, the main table matches first.

Policy routing priorities​

Linux evaluates policy routing rules in ascending priority order, and the first match wins. The relevant neighborhood on a Tailscale node looks like this:

PriorityRulePurpose
0locallocal addresses
5000local-network protectionkeeps LAN destinations in the main table
5210fwmarkTailscale's own marked traffic
5270table 52routes accepted from subnet routers
32766mainnormal routes
32767defaultfallback

The protection rule has to sit between the local rules and the accepted-routes rule. Priority 5000 is the conventional choice because it is high enough to beat 5270 but low enough to leave Tailscale's own marked traffic alone.

Best practices​

  • Do not advertise a subnet from a subnet router that lives inside that same subnet.
  • If the overlap is unavoidable, add the protection rule to every node on the subnet, not only to the one that failed.
  • Treat ICMP redirects from the router as a routing smell worth investigating.
  • Infrastructure devices that do not need remote subnet access should not accept routes at all.

Appendix: what disabling route acceptance does and does not do​

Turning off route acceptance stops a node from installing routes advertised by other subnet routers. It does not block inbound connections to that node.

The direction matters. If a device does not need to reach remote networks through the tailnet, disabling route acceptance removes the outbound table-52 path and is the simplest prevention. If the device must accept routes, the protection rule above keeps its replies on the LAN without giving up access to remote subnets. The two scenarios:

  • No route acceptance: no table-52 entries, no asymmetric path, no protection rule needed.
  • Route acceptance: table-52 entries exist; add the protection rule so LAN traffic still prefers the main table.

Takeaway​

When a host is unreachable from its own network but fine over Tailscale, suspect policy routing asymmetry before replacing hardware: an overlap between an advertised subnet and a local interface is enough to black-hole traffic while every status indicator stays green.

· 2 min read

The local network runs two DNS resolvers — a primary and a replica — kept in sync so that either can answer queries. That is the shape of high availability. What an audit found was that the replica had been offline for about a month, the primary had been answering everything alone, and nothing had said a word.

It kept going: the sync container's health check had failed tens of thousands of times in a row, and the webhook that was supposed to report sync failures had been returning 404 for longer still. Three independent problems, stacked so neatly that each one hid the next.

The layers that failed​

replica node down      → resolution still works (primary answers)
sync job failing → no visible symptom (replica not serving)
notification broken → failure report goes nowhere (webhook 404)

Any one of these being healthy would have surfaced the others. The redundancy worked so well that the failure was invisible.

What actual HA requires​

  • Monitor the replica, not the service. "DNS resolves" only proves that a resolver is up. A replica check has to ask the replica directly — a query against its own address, not the shared name.
  • Check the alert delivery path. A notification channel is a dependency like any other. Stale webhooks, expired tokens and changed URLs all fail silently unless something tests them. A periodic test alert or a dead man's switch turns "nobody was told" into a detectable condition.
  • A failing health check must reach a human. Repeatedly failing health checks that only appear in container logs are decoration.
  • Test failover. Until the replica has actually served queries while the primary was down, "HA" is a hope. A planned failover test is the only proof.
  • Make rejoining automatic. Firewall rules and reconnect configuration have to survive a reboot, or a restarted replica stays offline — which is exactly how a one-day outage becomes a month.

The uncomfortable takeaway​

The system did not fail because DNS stopped working. It failed because the redundancy was never exercised, and the reporting chain had the same blind spot as the thing it reported on. Redundancy without a test is just a second copy of the same assumption.

· 5 min read

An SDK's error behaviour is part of its public API. Every caller has to make a decision about errors — handle them, retry them, or ignore them — and they can only make that decision if the error they receive is honest about what happened.

Ours was not honest. The backend had two ways of reporting failure, and the client only handled one of them.

Two error channels​

GraphQL gives an API two distinct places to put an error.

Protocol errors live in the top-level errors array:

{
"errors": [
{
"message": "The current user is not authorized to access this resource.",
"extensions": { "code": "AUTH_NOT_AUTHENTICATED" }
}
],
"data": { "alert": null }
}

These are system-level failures: missing authentication, broken requests, legacy operations. Our HTTP client already converted them into a generic ClientError, so this channel worked.

Payload errors live inside the response data:

{
"errors": null,
"data": {
"createSlackChannelTarget": {
"slackChannelTarget": null,
"errors": [
{
"__typename": "TargetLimitExceededError",
"message": "You have reached the maximum number of targets."
}
]
}
}
}

This channel carries business-rule failures: quota limits, missing targets, invalid arguments. It is type-safe by design — the __typename tells you exactly which error you got.

And it was largely ignored. Mutation methods returned the payload to the caller and left the errors array for them to notice. Some did not even use the pattern yet. The result was an API surface where the same failure could be a rejection, a silent no-op, or a null field depending on which mutation you called.

What good looks like​

We wanted three properties:

  • One type to check. Callers should be able to ask "is this an authentication error?" without string matching a message.
  • No lost context. The original __typename, the message and the underlying cause should all survive the conversion.
  • No big-bang break. Existing throws had to keep working while new code adopted the new type.

The design​

We added an errors/ module to the client package rather than replacing the existing error handling. At its centre is a base class and a factory:

export class SdkError extends Error {
readonly errorType: string;
readonly code: string;
readonly cause?: unknown;

static from(e: unknown): SdkError {
if (e instanceof SdkError) return e;
if (isPayloadError(e)) return fromPayloadError(e);
if (e instanceof Error) return new SdkUnknownError(e.message, 'UNKNOWN', e);
return new SdkUnknownError(String(e), 'UNKNOWN', e);
}
}

The factory is the key detail. Instead of asking every call site to know which class to construct, they all funnel through SdkError.from(...). It is idempotent, it accepts anything, and it never throws while trying to describe a throw.

A type guard identifies payload errors by their shape:

function isPayloadError(
e: unknown,
): e is { __typename: string; message: string } {
return (
typeof e === 'object' &&
e !== null &&
'__typename' in e &&
'message' in e
);
}

Then a switch maps error families onto subclasses:

TargetLimitExceededError, TargetDoesNotExistError, Web3TargetNotFoundError, ...
→ SdkTargetError (TARGET)
UnauthorizedAccessError
→ SdkAuthenticationError (AUTHENTICATION)
ArgumentError, ArgumentOutOfRangeError, ArgumentNullError
→ SdkValidationError (VALIDATION)
everything else
→ SdkUnknownError (UNKNOWN)

Each subclass carries an errorType category, the original backend code, a timestamp and the cause. Consumers can now branch on category instead of probing messages.

At the call sites, mutations validate their payload before returning:

const mutation = await this.service.createWebPushTarget(input);
const errors = mutation.createWebPushTarget.errors;
if (errors && errors.length > 0) {
throw SdkError.from(errors[0]);
}
return mutation;

In the React layer, unsafe casts disappeared:

// before
.catch((e) => setError(e as Error))

// after
.catch((e) => setError(SdkError.from(e)))

The unglamorous part​

The interesting engineering was not the class hierarchy. It was the schema archaeology needed to make the hierarchy complete.

Payload errors only exist if the schema declares them, and it did not declare them consistently. We catalogued every mutation the SDK consumed, recorded which ones implemented the pattern, and then added the missing error fragments to the schema and type generation. Only after the types were honest could the runtime be.

A few operations also turned out to return protocol errors where the schema promised payload errors. The abstraction had to tolerate both — which the from() factory does by construction.

The deliberate breaking change​

The migration shipped in a major release, because it changed observable behaviour: mutations that previously swallowed payload errors now throw them. The release notes called it out explicitly, with a migration example:

try {
await client.deleteAlerts({ ids: alertIds });
} catch (error) {
if (error instanceof SdkValidationError) {
// handle validation error
}
}

That is the honest version. The previous behaviour — accepting an empty ID list and returning as if it did something — was the actual bug.

Why categories matter​

Unified errors are not just nicer to catch. They make automations possible.

We wanted to wire an on-call paging tool into the SDK's CI/CD pipeline, but paging a human on every expected failure is worse than no paging at all. Categorised errors let the pipeline ignore known conditions such as rate limits while still escalating genuine faults. Without the abstraction, that classification would have been string matching on messages — which breaks the first time a message is reworded.

If an SDK reports failures in two ways, callers will handle one of them and ignore the other. Collapsing both into a typed, categorised error is the smallest change that makes the API tell the truth.

· 4 min read

A monitoring stack usually gets installed for its dashboards and then trusted for its alerts. But an alert is a pipeline: a rule evaluates, a state changes, a notification is routed, a person is interrupted. Every stage can fail quietly. This is what that pipeline looks like in a small cluster and where the sharp edges are.

Two systems, two jobs​

  • Prometheus scrapes metrics, evaluates rules, and decides when an alert is firing.
  • Alertmanager receives firing alerts and decides how to group, route, inhibit and deliver them.

Keeping the split in mind prevents a common confusion: a rule that never fires and a notification that never arrives are different problems in different systems.

The life of an alert​

Inactive → Pending → Firing

A rule's expression becomes true, the alert enters Pending, and it stays there until it has been continuously true for the rule's for duration. Only then does it become Firing and get sent to Alertmanager. The for window is the difference between "this spiked for one scrape" and "this is broken".

Where the rules come from​

The kube-prometheus-stack chart ships rule packs that can be toggled on and off: application-level rules, node rules, and so on. One toggle is worth calling out — datastore rules are typically disabled by default because managed Kubernetes distributions own the datastore, and enabling the rules without the matching metrics only produces confusion.

To see what is actually installed:

kubectl get prometheusrules
kubectl get prometheusrule <name> -o yaml

The object YAML is the ground truth: the expression, the threshold and the for duration. Reading the live state is the Prometheus UI's Alerts page, which shows each rule as inactive, pending or firing with the current value.

A representative rule looks like this — a pod stuck in a crash loop:

alert: KubePodCrashLooping
expr: max_over_time(kube_pod_container_status_waiting_reason{reason="CrashLoopBackOff"}[5m]) >= 1
for: 15m
labels:
severity: warning

Know which metrics you actually have​

Two exporters carry very different information:

  • kube-state-metrics reports the state of API objects: how many replicas a deployment wants, whether a pod is waiting, and why.
  • cAdvisor reports container resource usage: CPU, memory, filesystem.

If cAdvisor is dropped to save resources — a reasonable choice on weak hardware — the object-level rules keep working, but every rule that depends on container memory or CPU silently has no data. Nothing breaks; the alert just never fires. Worth writing down at install time.

Grouping, timing and silencing​

Alertmanager's routing tree decides where alerts go; grouping decides how many messages a person receives. A crash-looping app with a bad replica count can produce several alerts at once, and group_by collapses them:

group_by: [namespace, alertname, severity]

Group too little and the phone buzzes per pod. Group too much and unrelated problems arrive glued together. The timing knobs are:

  • group_wait — how long to collect alerts before sending the first message.
  • group_interval — how often to send updates about an existing group.
  • repeat_interval — how often a still-firing alert is repeated.

Silences are the maintenance tool: a time-boxed mute that expires on its own. They are better than editing rules for planned work, but a silence that is too broad hides new problems inside its scope.

The blind spot: whitebox without blackbox​

Most of this stack is whitebox monitoring — it reports from inside the system. Pods say they are running, services say they exist. What it cannot tell you is whether a user can reach anything.

The classic gap: the ingress is broken, every pod is healthy, and Prometheus is happy. The cluster is "green" and the site is down.

Blackbox probing closes this: probe the important endpoints from outside on a schedule and alert when the response is wrong. If only one thing is added after the initial setup, this is the one with the highest return.

Recording rules​

On a small cluster, dashboards that recompute expensive expressions on every load can cost more than the monitoring is worth. Recording rules precompute an expression into a new metric:

record: job:request_rate:5m
expr: sum(rate(http_requests_total[5m])) by (job)

The dashboard then reads a cheap series. This matters most on low-power hardware, where a heavy query is competing with the workload it observes.

Takeaways​

  • A firing rule and a delivered notification are separate systems; design and test both.
  • Know which metrics are absent before trusting an alert to cover something.
  • Use for durations to filter noise, and group intentionally.
  • Whitebox monitoring cannot see broken entry paths. Add a blackbox probe.
  • On weak hardware, recording rules are a performance feature, not a luxury.

· 9 min read

SmartLink started with a product requirement that sounds simple until the SDK boundary becomes part of the design: let a partner define an interactive action in a backend admin experience, render it through the SDK, and allow an end user to execute the action from the host application.

Some of those actions could result in blockchain transactions.

I designed the SDK side of SmartLink around one constraint: the SDK could describe and initiate an action, but it should not own the customer's blockchain execution environment.

That decision shaped the client model, the package boundaries, the React API and the way transaction execution was handed back to the host.

The product shape​

The goal was to let configuration live outside the customer's application code.

A partner could define a SmartLink and its actions through backend tooling. The frontend SDK would fetch that configuration and render the corresponding inputs and actions.

At a high level:

This was important for product velocity. Adding or changing a configured action should not require every host application to hard-code a new UI or ship a new integration just to reflect backend configuration.

But it also created a harder question: what happens when an action needs to produce and execute an on-chain transaction?

Where I wanted the boundary​

It would have been easy to let SmartLink keep expanding until it owned the whole execution path:

That would make a demo work quickly, but it would also make SmartLink responsible for every wallet and chain decision made by the host application.

The host already knows which wallet system it uses. It already owns connection state, wallet UX, chain selection and transaction submission behavior. Importing those responsibilities into SmartLink would duplicate state machines and couple a backend-configured feature to an unstable wallet ecosystem.

So I kept the boundary narrower:

The SDK owns the product workflow. The host owns the execution environment.

That is the decision that keeps SmartLink blockchain-agnostic at its public boundary.

Backend configuration is a contract, not just JSON​

The configuration flow crosses several layers:

The backend configuration may arrive as serialized data, but I did not want arbitrary JSON leaking through the entire React tree. The SDK parses the response into a known SmartLink model and validates the shape before components depend on it.

That gives each layer a specific responsibility:

  • backend defines the product configuration
  • client interprets the configuration as SDK domain data
  • React context manages interactive state
  • components render the experience

This is the same principle I prefer elsewhere in SDK design: framework components should consume a domain model, not become the place where backend data is interpreted.

A separate client for a separate lifecycle​

One early design question was whether SmartLink should simply become another method set on NotifiFrontendClient.

I chose a sibling client instead.

The normal frontend client had accumulated a lifecycle around user authentication, persisted authorization and broader Notifi application state. SmartLink did not need to inherit that lifecycle just because it talked to some of the same services.

Its needs were narrower. It needed environment and configuration access, and some actions could use lightweight auth parameters such as a wallet public key and blockchain type. It did not need the main client's full persistent session model.

Conceptually:

Making SmartLink a subclass would have made reuse look elegant in the type hierarchy while coupling it to behavior it did not actually require.

Similar services do not imply the same lifecycle. That was the reason to prefer sibling clients over inheritance.

Why I did not create a completely isolated package​

Another possible boundary was a standalone SmartLink package containing its own client, React components, types and styling.

That sounds clean until the dependency graph is drawn.

SmartLink still needed infrastructure that already existed in the SDK:

  • shared GraphQL types and services
  • frontend domain models
  • React context conventions
  • theme and CSS-variable infrastructure
  • common error and loading behavior

A new package would either duplicate those systems or depend back on the existing frontend and React packages anyway. The isolation would be mostly organizational, not architectural.

Instead, I placed the feature into the layers where each responsibility already belonged:

The feature remained additive. Existing consumers did not need to import it, while the implementation could reuse the SDK's established infrastructure.

Reads and writes were intentionally different​

SmartLink has two service interactions that look related from the UI but have different operational characteristics.

I kept those dependencies explicit in the client rather than hiding them behind one generic service abstraction.

The distinction is useful because the configuration is tenant/link-level data that can be fetched and reused, while action execution is user-specific and may require current wallet identity or other runtime inputs.

An abstraction should remove accidental complexity, not erase meaningful differences between operations.

State is keyed by the domain, not by component instances​

A SmartLink can contain multiple actions, each with its own configured inputs and runtime values. A page can also render more than one SmartLink under the same provider.

I modeled the context state around domain identifiers rather than component-local state:

SmartLinkConfigDictionary
linkId → configuration

ActionDictionary
linkId:actionId → action state + user inputs

This did two things.

First, configuration could be fetched once and reused by multiple components instead of being tied to whichever component mounted first.

Second, input state had a stable identity even as individual input components mounted or unmounted. Later fixes around initialization and reset behavior reinforced the same idea: the action state is the source of truth; the input widget is only a view onto it.

That is a small implementation choice with a large effect on component reliability.

preAction instead of importing wallet state​

The host application sometimes needs something to happen before the configured action can execute. The most obvious example is wallet connection.

SmartLink could have added APIs like:

isWalletConnected
connectWallet
selectedChain
openWalletModal

I deliberately did not do that.

Instead, the React component exposes a small preAction extension point with behavior such as a label, disabled state and click handler. The host can use it to gate execution behind wallet connection or another prerequisite.

That means the SDK can render:

[ Connect wallet ]

when the host wants a prerequisite, and then render the normal action once the prerequisite is satisfied — without SmartLink learning anything about the host's wallet implementation.

This is an important design pattern for public SDKs:

When behavior varies by host application, expose a boundary instead of importing the host's state machine into the SDK.

Action execution stays chain-agnostic​

When the user submits an action, SmartLink validates the configured inputs and sends the action request to the backend/dataplane layer.

A simplified flow is:

The action request carries actionId, authParams and the configured input values.

The important part is what is missing from the SmartLink API: there is no EVM-specific signer contract, no Solana wallet adapter and no chain-specific transaction UI.

The backend and client can agree on an execution payload while the host decides how that payload becomes a real transaction in its environment.

That keeps the SDK API stable as wallet integrations evolve independently.

Hardening the design through real UI behavior​

The first version established the architecture, but SmartLink became a production feature through a sequence of smaller refinements:

  • action input validation and constraints
  • loading and inactive states
  • reset behavior
  • theme support
  • banner and tenant metadata
  • pre-action behavior
  • better context ownership of fetched configuration
  • unit and component coverage
  • Cypress coverage for success and failure paths
  • explicit handling of unmatched blockchain configuration

Those changes matter because configuration-driven UI has a lot of state that static component APIs do not: defaults, required fields, invalid values, temporarily unmounted inputs, inactive actions and backend-defined variations.

The architecture had to survive those cases without pushing special-case logic back into the host application.

The broader lesson​

SmartLink was not mainly a React component project. The React UI was the visible part of a boundary-design problem.

The design worked because responsibilities stayed separated:

The lessons I carried forward were:

  • Backend-driven configuration works best when the SDK interprets it into a real domain model instead of passing raw data through the UI.
  • Different lifecycle requirements deserve different clients even when they share services.
  • A package boundary is only useful when it reduces dependencies; a package that immediately depends back on the system it was meant to isolate is usually just indirection.
  • Host-specific prerequisites should be extension points, not SDK-owned state machines.
  • Blockchain-agnostic behavior comes from deciding where chain-specific execution stops, not from pretending chains are identical.

The most important design choice was the simplest one to describe: SmartLink knows how to configure and initiate the action. The application that embeds it remains in control of actually executing it.

· 7 min read

"Support another wallet" sounds like a small frontend task until the application supports more than one blockchain ecosystem.

The UI may only need another option in a selector, but underneath that button are different discovery mechanisms, address encodings, connection lifecycles and signing APIs. If those differences leak into application code, every new wallet makes the integration surface harder to reason about.

I ran into this while building reusable customer-facing applications around the Notifi SDK. The application needed wallet-based authentication, but different customers could require different chains and wallet ecosystems. The application itself was supposed to stay reusable.

That forced a boundary: wallet variation could not live in the page.

The application needed a stable contract​

From the application's point of view, wallet authentication is conceptually simple:

is the wallet available?
connect
get the user's key or address
sign a message
disconnect

The implementation is not simple.

An EVM wallet may expose an EIP-1193 provider and use a hex address. A Cosmos wallet may expose a bech32 address and return a different signature structure. Solana uses another key representation and transaction model. Cardano wallets introduce their own APIs and encoded address formats.

If the host application owns those details, authentication quickly becomes a tree of chain-specific branches.

The reusable application would then need to know both product behavior and wallet protocol behavior:

That is the wrong dependency direction.

Move ecosystem knowledge behind a provider boundary​

The answer became a dedicated package: @notifi-network/notifi-wallet-provider.

The package owns the chain- and wallet-specific integrations and exposes a shared React-facing surface to the rest of the application.

The important part is not the React context itself. The important part is what the context prevents the application from needing to know.

At a high level, the boundary looks like this:

The public package now supports wallets across EVM, Cosmos, Solana and Cardano families. The application consumes one provider instead of importing each wallet implementation directly.

Normalize capabilities, not implementations​

A useful abstraction does not pretend all wallets are internally identical.

Instead, it defines the capabilities the application actually needs.

The public wallet types converge on concepts such as:

isInstalled
walletKeys
connect
disconnect
signArbitrary
sendTransaction // where supported
websiteURL

Behind that contract, wallet-specific implementations are still free to differ.

For example, key material is normalized into a structure that can represent several encodings:

hex
bech32
base58
base64
cbor

The application does not have to collapse those formats into one fictional universal address. It asks the selected wallet for the representation needed by the authentication path.

That is an important distinction. Good abstraction removes irrelevant variation; it does not erase meaningful protocol differences.

Signing is where the differences become real​

Message signing is a good example of why a shared interface helps.

An EVM wallet can sign a UTF-8 message and return a hex signature. A Cosmos wallet may accept bytes and return a structured signature. Solana and Cardano expose different APIs again.

The integration layer still needs to turn those outputs into the shape expected by the SDK authentication flow, but it can do so through one selected-wallet path instead of embedding each wallet library throughout the application.

The application becomes responsible for adapting one normalized wallet object to the SDK auth contract.

The wallet provider becomes responsible for adapting many external wallet protocols to the normalized wallet object.

Those are much cleaner responsibilities.

A factory boundary makes new wallets cheaper​

Once wallet behavior is behind one interface, adding support becomes a local change instead of an application-wide one.

Conceptually:

That does not mean every wallet is zero-cost. Detection can be inconsistent. Some providers initialize late. Standards evolve. Hardware wallets and chain-specific signing flows create special cases.

But the cost is contained.

The customer application does not need a new routing model, a new auth context or a second copy of the subscription UI because another wallet was added.

The abstraction also improved product delivery​

This wallet layer was not created in isolation. It supported a broader delivery model built around the public notifi-dapp-example.

That application is intended to be reused across integrations. A customer's branding and configuration may change, and so may the required chain or wallet. If those dimensions were coupled, each combination would become another bespoke application:

Customer A + EVM + MetaMask
Customer B + Solana + Phantom
Customer C + Cosmos + Keplr
Customer D + Cardano + Lace

With a wallet boundary, the shape is different:

The customer-specific choice moves into configuration and supported-provider selection rather than application architecture.

Standards still matter inside an abstraction​

A provider layer should not become an excuse to permanently hide outdated integrations.

Wallet ecosystems change. Browser wallet discovery standards improve, packages get replaced and providers change how they inject themselves into the page. The abstraction gives us one place to respond to those changes, but the adapter itself still needs to follow the ecosystem.

That was useful when wallet discovery moved toward standards such as EIP-6963. The host application did not need to learn a new discovery strategy; the wallet layer could evolve behind the same product-facing contract.

This is one of the biggest benefits of an adapter boundary around third-party ecosystems: change remains inevitable, but its blast radius becomes deliberate.

Where I would draw the boundary again​

If I were designing the same system from scratch, I would keep the same basic rule:

The application should understand authentication requirements, but it should not understand wallet implementation details.

I would also make three things explicit early:

Capability contracts. Define the smallest set of operations the product needs instead of exposing raw provider objects everywhere.

Key formats. Treat address and public-key representation as part of the chain contract. Do not normalize away information the backend actually needs.

Registration. Make wallet support declarative enough that adding a provider mostly means implementing an adapter and registering its capabilities.

Those constraints make the package easier to extend without turning the abstraction into another monolith.

The broader lesson​

The wallet provider solved a wallet problem, but the architectural lesson is more general.

Reusable customer software works best when unstable external ecosystems sit behind narrow boundaries. The stable product should depend on the capability it needs, not on every vendor-specific way that capability can be delivered.

For this system, the stable requirement was simple: connect an identity and produce the signing behavior needed for authentication.

Everything below that line could change — wallet vendor, chain, discovery mechanism, address encoding, signing API — without forcing the customer application to become a different product.

· 6 min read

An SDK is a good delivery mechanism when the customer already has an application and wants to embed your product into it. It is a much less complete answer when the customer wants you to deliver the application too.

That distinction changed how I thought about the frontend surface around our SDK.

At Notifi, the original integration model was straightforward: we shipped client and React SDKs, and customers embedded notification functionality into their own web applications. The SDK owned the product logic; the customer owned the surrounding experience.

Then a different requirement started appearing. Some customers did not want another component to integrate. They wanted a complete page they could send their users to: connect a wallet, authenticate, configure notification subscriptions and talk to the same Notifi backend without first building an application around the SDK.

The obvious implementation was also the wrong long-term one: build a new app for every customer.

The repeated work was the signal​

A customer-specific full-page integration has obvious differences:

  • branding and visual style
  • tenant credentials
  • notification configuration
  • blockchain and wallet expectations
  • page copy

But most of the application does not change:

  • authentication flow
  • SDK/provider wiring
  • subscription state
  • notification configuration UI
  • routing
  • loading and error handling
  • deployment shape

If every customer starts from an empty repository, the stable 80 percent gets rewritten because the variable 20 percent looks different.

That is not customization. It is duplication with a different logo.

The boundary I wanted was the opposite: keep the product flow stable and make customer-specific variation explicit.

A full-page example became the reusable baseline​

The result was notifi-dapp-example, a Next.js application inside the public SDK monorepo.

The name says "example", but its role grew beyond documentation. It serves two related use cases:

  1. a runnable reference for developers learning how the SDK pieces fit together
  2. a baseline application that can be cloned, branded and configured for a customer deployment

That dual purpose matters. A conventional example usually optimizes for readability and throws away production concerns. A customer template has to survive them: environment configuration, real authentication, error states, deployment and continued SDK evolution.

Instead of forking core business logic, the application consumes the public SDK packages just like an external integration would. That keeps the template honest. If the public integration surface becomes difficult to use, the example feels the pain too.

Put variation into configuration​

The next question was deciding what should require code changes.

The application exposes customer-specific values through configuration rather than scattering them through components. The public example includes variables for things such as:

tenant
runtime environment
blockchain
subscription card
page title
page subtitle

That is a small design decision with a large operational effect.

For a normal customer variation, the workflow becomes roughly:

The application architecture is no longer part of every customization project.

This also gives customers a useful escape hatch. If they want to host the experience themselves, the same public application shows a complete Next.js implementation rather than only a collection of SDK snippets. They can inspect it, run it and adapt it without depending on an internal codebase.

Deployment is part of the template​

A reusable application is not reusable if every fork needs a new deployment design.

The public package therefore also documents a deployment path built around development and production environments. The application configuration is passed through environment variables, and the deployment flow can build the app and publish the resulting assets through the same repeatable pipeline.

This is an important distinction between a sample and a delivery baseline:

Once deployment is repeatable, customer work moves away from infrastructure invention and toward the places where customization is actually valuable.

Wallets were the remaining source of structural variation​

Branding and credentials are easy to parameterize. Wallet authentication is harder.

A customer on an EVM chain may expect MetaMask or WalletConnect. A Solana integration may need Phantom. A Cosmos integration may use Keplr. Other ecosystems have different wallet APIs, address formats and signing behavior.

If the template handled those differences directly, every new wallet would spread another branch through the application:

if EVM ...
if Solana ...
if Cosmos ...
if Cardano ...

At that point the reusable shell stops being reusable.

So wallet variability became a separate package boundary: @notifi-network/notifi-wallet-provider.

The full-page application consumes a normalized wallet layer; the provider package absorbs wallet- and chain-specific behavior. A new customer can therefore change the supported wallet set without replacing the application's authentication architecture.

That separation turned out to be one of the most important parts of making the template durable.

The architecture became a delivery system​

The useful mental model is not "SDK plus example app." It is a set of integration levels built on the same product surface:

The SDK remains the source of product behavior. The full-page app provides a stable integration shell. The wallet layer isolates one of the messiest external dependencies. Configuration carries the customer-specific values.

No layer needs to pretend every customer is identical, but customer differences no longer force us to redesign the entire system.

What I learned​

The main lesson was that reusable software is less about identifying identical code than identifying where variation is allowed to enter the system.

A template with dozens of customer-specific conditionals is only a shared repository. A reusable integration platform has a stable center and explicit extension points.

For this system, those extension points became:

  • configuration for tenant and product behavior
  • styling for customer identity
  • a wallet-provider boundary for chain-specific authentication
  • public SDK APIs for the underlying notification product

That made the same codebase useful as documentation, as a starting point for self-hosting and as the baseline for repeated customer delivery.

The interesting part was never the fact that we had an example application. It was turning repeated implementation work into an architecture where the next integration was mostly configuration instead of reinvention.

· 8 min read

We did not start with a framework-agnostic client. We started with React because that was where our customers were.

Early on, the product goal was simple: make it as fast as possible for a customer to add Notifi to an existing web application. Most of those applications were built with React, so a React-first SDK was the shortest path to a useful integration. We packaged service access, authentication and subscription behavior behind hooks and paired them with a component library that customers could drop into their applications.

That trade-off worked. It reduced integration work and helped us get the product in front of customers quickly.

The architecture only became a problem after the customer base grew.

The architecture we optimized for first​

The early stack looked roughly like this:

The hooks were not just React bindings. Over time they accumulated responsibilities that belonged to the product domain:

  • authentication and token state
  • targets and target groups
  • alert creation and subscription behavior
  • notification history
  • configuration fetching and interpretation
  • chain-specific signing behavior
  • orchestration across service calls

For a React application this was convenient. The component could call a hook, the hook knew how to perform the operation, and the customer did not need to understand the underlying service model.

The problem was that the business logic and the framework boundary had become the same thing.

Success changed the constraint​

As integrations expanded, React was no longer a safe assumption. Some customers used Vue. Others used Angular. Some wanted to integrate the service without adopting our React component library at all.

At that point the existing API had an architectural limitation: even when the operation itself had nothing to do with React, consuming it meant pulling in a React-specific layer.

The problem was not that hooks were inherently wrong. They were the right optimization for the first customer set. The problem was that they had become the owner of domain behavior that other frameworks also needed.

Splitting one large hook into more hooks would not solve that. useAuth, useAlerts and useTargets would be cleaner React code, but the product model would still be trapped behind React lifecycle and context.

The boundary needed to move.

The target architecture​

The consolidation plan was to make notifi-frontend-client the owner of frontend domain behavior and let React consume it like any other client:

The key distinction was responsibility.

NotifiFrontendClient would own things that should behave the same regardless of UI framework:

authentication state
configuration interpretation
alerts and subscriptions
targets and target groups
notification history
storage
service orchestration
chain-specific domain behavior

Framework packages would own the things that actually are framework-specific:

rendering
context / providers
component state
framework lifecycle
UI composition

This was more than converting hooks into class methods. It was turning React from the location of the SDK's business logic into one consumer of a stable domain API.

Building the client before replacing the hooks​

The new client was introduced gradually rather than as a rewrite.

The first notifi-frontend-client work established a TypeScript API facade over the GraphQL service. Its intended direction was already broader than React: business logic should be reusable across frontend frameworks and, where possible, across blockchain integrations.

The next step was capability parity. Before the React stack could depend on the client, the client had to cover the behavior that the hooks already provided.

That meant moving or consolidating operations such as:

  • initialization and persisted auth state
  • login and logout
  • target-group operations
  • alert creation and deletion
  • subscription-card configuration
  • notification history
  • wallet-related subscription behavior
  • conversation operations used by support UI

It also meant expanding the client across the event types and chains that the existing React packages already supported. A framework-agnostic abstraction is not useful if customers still have to fall back to a React hook whenever they hit an older feature.

Configuration became a domain concern​

One important part of the extraction was configuration.

The backend already stored tenant-level configuration for embeddable experiences. Instead of making each React component understand the raw backend shape, NotifiFrontendClient became the layer that fetched and interpreted that configuration.

A simplified flow became:

That separation matters because backend-driven UI is much easier to evolve when the rendering framework is not also responsible for interpreting the product domain.

The backend can own what is configured. The client can own what that configuration means operationally. React can own how it is presented.

This also reduced duplicated models. During the migration, React components increasingly consumed the types and configuration models exposed by the frontend client instead of maintaining parallel representations.

Migrating a live SDK without a big-bang cutover​

The most important design decision was not the client class itself. It was the migration path.

Existing customer integrations already depended on the hooks and React card packages. Replacing their implementation in one release would have made every behavioral mismatch a customer-facing regression.

So the migration moved in independently shippable stages:

During 2023 the React card progressively gained a frontendClient path for fetching data, rendering subscription state and executing individual event-type operations. For a period, components could run either implementation.

That duplication was intentional. It created a compatibility window where we could compare behavior and fix gaps without forcing every customer onto the new architecture at once.

Later that year, the default flipped: React used FrontendClient unless an integration explicitly chose the legacy hooks path.

That was the real migration milestone. The new client was no longer an experiment running next to the SDK; it had become the SDK's default domain implementation while the old path remained available as a safety valve.

The awkward middle was part of the design​

Running two implementations introduced its own problems.

Initialization order mattered. React could not safely render children before the frontend client had restored its state. There were also cases where both the hook path and the frontend-client path could trigger rendering work, creating race conditions or duplicate updates.

Those issues are easy to interpret as evidence that a migration should have been done all at once. I see them differently. They were the cost of preserving compatibility while changing a public SDK underneath live integrations.

The important part was keeping that period temporary and directional: every new piece of business logic moved toward the client, not back into the hooks.

Completing the transition​

The framework-agnostic client eventually became the foundation for the newer notifi-react package as well as other integrations. By 2024 the old stack — including notifi-react-hooks, the legacy React card, the old core and Axios adapter packages — could be deprecated and removed from the workspace.

The resulting architecture was much simpler conceptually:

React still had a first-class integration. It just no longer defined the SDK's domain architecture.

That distinction became increasingly valuable as the product expanded. New authentication flows, target types, wallet behavior and backend configuration could evolve in the client without requiring the domain implementation to be rewritten around a React hook.

What I took away​

Optimize for the market you have, but know which decisions are temporary. Starting with React was the right time-to-market decision. Treating React as the permanent owner of business logic would not have been.

Framework APIs should orchestrate UI, not own domain behavior. Hooks are a good ergonomic surface. They are a poor portability boundary when every important operation only exists inside them.

A client facade should model product concepts, not just wrap HTTP calls. The value of NotifiFrontendClient came from giving authentication, subscriptions, targets and configuration a stable API independent of React and transport details.

Migration compatibility is part of architecture. Supporting the old and new implementations side by side was not elegant, but it made the boundary movable without turning the refactor into a coordinated customer migration.

The clean architecture is usually the end state, not the starting point. The useful question is not whether the first version was perfectly decoupled. It is whether the system can evolve when the assumptions that made the first version successful stop being true.