Skip to content

[SDK-673] Add switchProject for runtime project switching - #1086

Draft
joaodordio wants to merge 1 commit into
feature/SDK-675-offline-disable-devicefrom
feature/SDK-673-switch-project
Draft

[SDK-673] Add switchProject for runtime project switching#1086
joaodordio wants to merge 1 commit into
feature/SDK-675-offline-disable-devicefrom
feature/SDK-673-switch-project

Conversation

@joaodordio

@joaodordio joaodordio commented Aug 25, 2026

Copy link
Copy Markdown
Member

📝 Summary

Adds IterableApi.switchProject(context, apiKey, config, callback), moving a running app between Iterable projects in place with no restart and no state carried over.

🎟️ Jira Ticket: SDK-673

⚠️ Stacked PR. This targets feature/SDK-675-offline-disable-device, not master. Review #1085 first and merge it first. SDK-673 depends on it: the "queued device disables survive the switch" behaviour is inert without the offlineApiSet change in SDK-675, because a disableDevice is not offline-queueable without it. Re-target this to master once SDK-675 lands.

The iOS counterpart is iterable-swift-sdk#1091 (SDK-674). The two were written together and the behaviour is intended to match.

📖 Description

For multi-region apps that need to move between a US and an EU project without restarting. Previously the only reliable way to change projects was an app restart.

The call returns immediately and runs the whole sequence on the background executor: disable the push token on the previous project, reset the in-app, embedded and unknown-user managers, purge the offline queue apart from queued device disables, clear identity and the rest of the previous project's storage, then re-initialize against the new key. The auth manager is rebuilt against the new config and the request processor re-bound to it, which iOS gets for free from its instance swap.

Callback contract

IterableProjectSwitchCallback is a single-method interface, onProjectSwitched(boolean cleanTeardown), delivered on the main thread.

true means every teardown step completed cleanly. false means the SDK is on the new project but a cleanup step was noisy, or no device disable could be confirmed. false never means the switch failed or was rolled back. An app that does not use push always sees false, which is normal and not an error. The right response is the same either way: carry on and re-identify the user.

It is a single-method interface deliberately. Reusing IterableInitializationCallback does not work: its only abstract method takes no arguments, so a lambda would bind to that one and silently discard the boolean.

Details worth a reviewer's attention

The previous project's device disable. The disable captures that project's API key and its region endpoint when it is initiated, because the FCM token lookup is asynchronous and the live key can change underneath it. Without the captured key the disable lands on the new project, leaving the old project still delivering push to the device. Without the captured endpoint a cross-region switch sends the old key to the new region and is rejected. The switch waits up to 2 seconds for the disable to reach the request layer before swapping; if it times out the switch still completes and reports false, and the disable still reaches the project it was created for.

Region binding for offline tasks. Offline tasks now persist the endpoint they were created for, so a rehydrated request goes to the region it was built for rather than whichever region is live at flush time. Tasks already on disk from an earlier version keep resolving the old way, so nothing queued is lost on upgrade.

trackPushOpen is not queued behind the switch gate. A push payload carries the sending project's campaignId, templateId and messageId, so replaying it after the switch would report it against a project where those IDs do not exist. It runs inline instead. This needed care because the switch gate and the background-init gate are the same state on Android, so queueOrExecuteUnlessSwitching splits them: still queued during initialization, inline during a switch. iOS does not gate push handling at all, for the same reason.

Gate consistency. The longest track and updateEmail overloads were public and ran inline while every shorter overload was queued, so mid-switch behaviour depended on which overload the caller used. Both now wrap private *Internal methods, matching the existing setEmail / setUserId treatment.

Per-project state on the shared instance. iOS drops this when it replaces its SDK instance; Android reuses sharedInstance, so it has to be explicit. The switch now clears the inbox session ID, stored push payload, notification data and device attributes. The inbox session ID was the one producing cross-project data: it would otherwise be attached to the new project's first in-app tracking call. Device ID and visitor consent are project-agnostic and deliberately kept.

Concurrency. The gate check and enqueue are a single atomic step, so a call cannot pass the check just before the gate is raised and then run against a half torn-down SDK. Twelve fields are now volatile and getAuthManager()'s rebuild is guarded by a lock. The switch also recovers if the background executor is shut down when the teardown or drain is submitted, which could happen when switchProject was called from inside a switch callback.

Guards. A null context or apiKey throws IllegalArgumentException, since both are @NonNull and a null is a programmer error. An empty or whitespace-only key is a runtime condition, so it is refused without tearing anything down and reported as false, matching iOS.

Known limitations, called out deliberately

  • Campaign-attributed events other than push opens are still queued. trackPurchase with a campaignId, and track with a campaignId, are replayed against the new project carrying the previous project's IDs if they are issued during the switch window. iOS queues its campaign-carrying trackPurchase too, so this is consistent across platforms rather than an Android-only gap. Fixing it properly means per-call key binding on both SDKs and belongs in its own ticket.
  • iOS has no in-flight-initialize guard, Android does. On iOS an initialize immediately followed by switchProject is torn down underneath. That divergence needs an iOS follow-up.
  • Neither SDK persists switch intent, so a process death mid switch leaves the app restarted with the previous project's identity cleared and no record a switch was attempted.

🧪 How to test?

758 tests, 0 failures, checkstyle clean.

  • IterableSwitchProjectTest, 41 tests: the guard cases, each teardown step, identity and storage clearing, manager rebuild, keychain rebuild, offline queue region binding, rapid and nested switches, a throwing teardown step, callback delivery and thread, blank keys, per-project instance state, and gate consistency for the longest overloads.
  • IterableSwitchProjectQueueDrainTest, 4 tests: a drain no executor will accept, a switch started off the main thread, and the two push-open cases (inline during a switch, still queued during initialization).
  • IterableSwitchProjectDisableRegionTest: the disable dispatch timeout path, driven with an overridden timeout so it does not burn real time.
  • IterableOfflineTaskRegionTest, 13 tests: persisted baseUrl, rehydrated task region, old-schema fallback, cross-region isolation, disable preservation.
  • IterablePushRegistrationTaskTest: the disable carries the captured key and endpoint rather than the live ones, and is sent before the key swap.

Every new test was checked to fail with its fix reverted, not just to pass alongside it.

Manual check: initialize against project A, identify a user, then call switchProject with project B's key from the callback of a region lookup. Confirm the device is disabled on A and registered on B, the inbox is empty immediately after, and no event reaches A after the callback.

🧾 Changelog

Added to CHANGELOG.md under Unreleased. One Added entry for switchProject with sub-bullets covering the callback contract, the queued and inline call sets, the state that is cleared and preserved, and the edge cases. Seven Fixed entries covering offline endpoint persistence, disable preservation across a switch, the overload gating, the in-flight-initialize deferral, executor recovery, the duplicate auth listener, and the atomic gate check.

📹 Loom recording if applicable

Not recorded.

🐞 Github Issues solved

None known.

📚 Docs PR if applicable

A docs PR is required and does not exist yet. This adds public API. There is an adoption guide written for the first customer team, which should be the basis for the iterable-docs entry, and it needs to cover the callback contract, the per-platform queued call sets, and the iOS in-flight-initialize caveat.


Note on the base: this stack sits on 5d726694 and origin/master has moved 4 commits ahead, including a 3.10.1 release prep and SDK-547, which touches JWT auth timing. Worth rebasing both branches onto latest master before merge, since SDK-547 is adjacent to the auth manager rebuild here.

Adds IterableApi.switchProject(context, apiKey, config, callback), which moves a
running app from one Iterable project to another in place, with no app restart and
no state from the previous project leaking into the new one. Aimed at multi-region
apps that need to move between a US and an EU project without restarting.

The call returns immediately and runs the whole sequence on the background
executor: disable the push token on the previous project, reset the in-app,
embedded and unknown-user managers, purge the offline queue apart from queued
device disables, clear identity and the rest of the previous project's storage,
then re-initialize against the new key. The auth manager is rebuilt against the new
config and the request processor re-bound to it, which iOS gets for free from its
instance swap.

Callback contract: IterableProjectSwitchCallback is a single-method interface so a
lambda receives the result. true means every teardown step completed cleanly, false
means the SDK is on the new project but a cleanup step was noisy or no device
disable could be confirmed. false never means the switch failed or was rolled back.
An app that does not use push always sees false, which is not an error.

Notable details:

- The previous project's disable captures that project's API key and its region
  endpoint when it is initiated, because the FCM token lookup is asynchronous and
  the live key can change underneath it. Without the captured key the disable lands
  on the new project; without the endpoint a cross-region switch sends the old key
  to the new region and is rejected. The switch waits up to 2 seconds for the
  disable to reach the request layer before swapping.
- Offline tasks now persist the endpoint they were created for, so a rehydrated
  request goes to the region it was built for instead of whichever region is live
  at flush time. Tasks already on disk keep resolving the old way.
- trackPushOpen is not queued behind the switch gate. A push payload carries the
  sending project's campaignId, templateId and messageId, so replaying it would
  report it against a project where those IDs do not exist. It runs inline instead.
  Initialization queueing is unaffected.
- The longest track and updateEmail overloads are now queued like their shorter
  siblings. They were public and ran inline, so mid-switch behaviour depended on
  which overload the caller happened to use.
- Per-project state held on the shared instance is cleared: inbox session ID, push
  payload, notification data and device attributes. iOS drops all of this when it
  replaces its instance; Android reuses sharedInstance so it has to be explicit.
- The gate check and enqueue are a single atomic step, so a call cannot pass the
  check just before the gate is raised and then run against a half torn-down SDK.
- Recovers if the background executor is shut down when the teardown or drain is
  submitted, which could happen when switchProject was called from inside a switch
  callback.

checkstyle: FileLength stays suppressed for IterableApi.java only, tracked in
SDK-677.
final List<IterableInitializationCallback> initCallbacksToNotify = new ArrayList<>();
final ExecutorService executor;
synchronized (initLock) {
isSwitchingProject = false;

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The switching flags are currently cleared before the asynchronous drain starts, which could potentially allow newer calls to run before calls queued during the switch.

Consider keeping the gate raised until all queued calls have finished running.

Related comment on the iOS PR


// Step 2: raise the switch gate synchronously, so calls made after this method returns are
// queued rather than executed against a half torn-down SDK.
if (!IterableBackgroundInitializer.beginProjectSwitch(callback)) {

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

A second switchProject call with a different API key is currently ignored, although its callback still runs.
For example, if ProjectA -> ProjectB is running and the app requests ProjectA -> ProjectC, both callbacks fire but the SDK ends on B.

This may be surprising behavior for a multi-region API, but it also seems to be intentional and covered by tests, so this comment might be more of a design discussion than a blocker.

Should we clarify or reconsider the behavior when a second call targets a different API key?
For example, only joining requests targeting the same project and either queueing or rejecting requests targeting a different one?

Related comment on the iOS PR

// than being a key of its own, but remove it too so a build that starts writing it
// separately cannot carry a previous project's criteria id across a switch.
editor.remove(IterableConstants.SHARED_PREFS_CRITERIA_ID);
editor.remove(IterableConstants.SHARED_PREFS_ATTRIBUTION_INFO_KEY + IterableConstants.SHARED_PREFS_OBJECT_SUFFIX);

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

A deep-link redirect started on the previous project can finish after the switch and write its attribution into the new project’s shared storage.

Maybe the redirect could capture the project / API key it belongs to and discard the result if that project is no longer active?

Related comment on the iOS PR

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants