Migration¶
Move content between Microsoft 365 and other systems — resumable, verifiable,
and with best-effort fidelity. The toolkit is product-agnostic (a DataSource
→ DataTarget pair) and ships adapters for the filesystem, SharePoint lists,
and SharePoint document libraries.
The model¶
| Piece | Role |
|---|---|
DataSource / DataTarget |
adapters — read/checksum and exists/write/list_paths |
MigrationJob |
lifecycle assess → plan → run → verify, resumable |
Manifest |
the persisted plan (a list of MigrationItem) |
Checkpoint |
per-item status, incremental watermarks, phase (pause/resume/cancel) |
MigrationServerJob |
server-side Migration API ingestion (submit + monitor) |
Quick start¶
Filesystem → SharePoint document library:
from office365.migration import MigrationJob, MigrationOptions
from office365.migration.adapters.filesystem import FileSystemSource
from office365.migration.sharepoint.adapters import SharePointLibraryTarget
from office365.sharepoint.client_context import ClientContext
ctx = ClientContext(site_url).with_client_certificate(tenant, client_id, thumbprint, cert_path)
library = ctx.web.lists.get_by_title("Documents").root_folder
job = MigrationJob(
FileSystemSource("src"),
SharePointLibraryTarget(library),
MigrationOptions(concurrency=4),
)
job.plan()
job.run()
print(job.stats.summary())
print(job.verify().summary())
job.export_reports("reports/")
Runs are idempotent and resumable: pass manifest_path / checkpoint_path
and a paused or failed run continues where it stopped (job.resume()). Set
MigrationOptions(incremental=True) to skip items at or below the persisted
watermark.
Adapters¶
| Adapter | Direction | Content |
|---|---|---|
FileSystemSource / FileSystemTarget |
both | files + folders |
JsonFileSource / JsonFileTarget |
both | record payloads (one JSON file per item) |
SharePointLibrarySource / SharePointLibraryTarget |
both | files + folders (recursive) |
SharePointListSource / SharePointListTarget |
both | list items as records |
TeamsArchiveSource / TeamsArchiveTarget |
both | Teams archives |
SPMT-style sessions¶
MigrationSession mirrors the
SPMT PowerShell cmdlets
— register a session, add tasks, start, show, stop:
| cmdlet | method |
|---|---|
Register-SPMTMigration |
session.register(settings=..., context=...) |
Add-SPMTTask |
session.add_task(file_share_source=..., target_site_url=..., target_list=...) |
Show-SPMTMigration |
session.show() |
Start-SPMTMigration |
session.start() |
Stop-SPMTMigration |
session.stop() |
Unregister-SPMTMigration |
session.unregister() |
from office365.migration import MigrationSession, MigrationSettings
session = MigrationSession().register(
context=ctx, # a SharePoint ClientContext
settings=MigrationSettings(use_migration_api=True), # server-side Migration API
)
session.add_task(
file_share_source="C:/src",
target_site_url="https://contoso.sharepoint.com/sites/team",
target_list="Documents",
)
session.start()
print(session.show())
session.unregister()
MigrationSettings mirrors Register-SPMTMigration (permissions, versions,
filters, user mapping, Azure storage…); the subset the core enforces is mapped by
to_options(). Tasks can also be declared as MigrationTask objects or parsed
from the SPMT JSON task format (MigrationTask.from_json(...)).
Fidelity¶
| Flag | Client-side runner | Server-side Migration API |
|---|---|---|
preserve_timestamps |
best-effort (Created/Modified) |
full |
preserve_permissions |
best-effort (same-tenant ACL copy) | full |
preserve_versions |
✗ (raises) | full |
REST cannot restore version history, so preserve_versions always requires the
server-side API. The other two are applied client-side when the adapters
support them — the runner fails fast otherwise, rather than silently migrating
without the requested fidelity:
- Timestamps — the SharePoint library target restores
Created/ModifiedviaValidateUpdateListItem(the same path asListItem.system_update). - Permissions — the source reads the item's
role_assignments; the target breaks inheritance and recreates them, resolving principals by login name and roles by name. This is a same-tenant copy; cross-tenant identity mapping is not implemented.
Adapters opt in by implementing optional hooks:
DataSource.read_permissions(item), DataTarget.apply_timestamps(item), and
DataTarget.apply_permissions(item, permissions).
Server-side migration (full fidelity)¶
For version history and true ACL fidelity, content is packaged and ingested
server-side. The package layer builds the Migration API manifest XML
(Manifest.xml / ExportSettings.xml / SystemData.xml / UserGroup.xml)
from a document library's files and folders:
from office365.migration import MigrationJob
from office365.migration.adapters.filesystem import FileSystemSource
from office365.migration.sharepoint.package_target import SharePointPackageTarget
# Provision SharePoint-owned Azure containers (SAS URIs + encryption key)
containers = ctx.site.provision_migration_containers().execute_query().value
target = SharePointPackageTarget(
ctx.site,
web_id,
content_uri=containers.DataContainerUri,
manifest_uri=containers.MetadataContainerUri,
encryption_key=containers.EncryptionKey,
)
job = MigrationJob(FileSystemSource("src"), target)
job.plan()
job.run() # stages the package and submits the ingestion job
target.monitor() # polls GetMigrationJobProgress
MigrationServerJob can also be driven directly (submit / submit_encrypted /
progress / status_fn / monitor).
With encryption_key set, the staging AES-256-CBC encrypts every content and
manifest blob (unique random IV, stored as the base64 IV blob property) — which
SharePoint-provided containers require. For your own (BYO) containers, omit it and
use Site.create_migration_job instead of the encrypted variant.
The builder covers the document-library subset — webs, lists, folders, files,
and file versions. The generated XML follows the documented format but is not
yet verified against a live tenant (see office365.migration.package).
Storage & vendor neutrality¶
Server-side ingestion is the only leg that needs Azure — and that is Microsoft's requirement, not ours. The Migration API reads content/manifest blobs from Azure Blob Storage; it cannot read from local disk or S3.
| Move | Path | Azure? | Fidelity |
|---|---|---|---|
| M365 → filesystem / S3 / other | client-side adapters | no | content + best-effort timestamps/ACLs |
| filesystem / other → M365 | client-side adapters | no | content + best-effort |
| filesystem / other → M365 | server-side Migration API | yes | versions, ACLs, authors |
| M365 → M365 | either | only for server-side | — |
You don't need an Azure account¶
Site.provision_migration_containers() provisions SharePoint-owned containers —
no extra cost, and without the need to manually set up in the Azure admin
console. It returns the container URIs (with SAS tokens) and the encryption key.
The [azure] extra is only the client library used to push blobs into
whatever container you were handed — it is not an Azure subscription:
Step by step¶
from office365.migration import MigrationJob
from office365.migration.adapters.filesystem import FileSystemSource
from office365.migration.sharepoint.package_target import SharePointPackageTarget
# 1. Provision SharePoint-owned containers (no Azure account needed)
containers = ctx.site.provision_migration_containers().execute_query().value
# containers.DataContainerUri -> content container (SAS)
# containers.MetadataContainerUri -> manifest container (SAS)
# containers.EncryptionKey -> AES256CBC key
# 2. Point the target at the containers (staging is swappable — see below)
target = SharePointPackageTarget(
ctx.site,
web_id,
content_uri=containers.DataContainerUri,
manifest_uri=containers.MetadataContainerUri,
encryption_key=containers.EncryptionKey,
)
# 3. Migrate: items -> manifest XML -> staged blobs -> submitted job
job = MigrationJob(FileSystemSource("src"), target)
job.plan()
job.run()
print(target.job_id)
# 4. Poll GetMigrationJobProgress until the job is terminal
target.monitor()
The Staging seam¶
The package (documented XML + content blobs) is vendor-neutral bytes; where they
land is a strategy. Pass staging= to SharePointPackageTarget:
FileSystemStaging— writesmanifest/+content/to disk (no Azure; tests, inspection, or handing the package to another uploader);BlobStaging— Azure Blob (whatcreate_staging(...)returns today);create_staging(content_url, manifest_url)— the factory that picks a backend by container host. It is the extension point: a future S3 / Azure Files staging slots in here without changing callers.
from office365.migration.package import FileSystemStaging
# Build + write the package to disk (no Azure, no submit)
target = SharePointPackageTarget(
ctx.site, web_id, staging=FileSystemStaging("pkg/content", "pkg/manifest")
)
package = target.stage()
When you don't need Azure at all¶
If you don't need version history or ACLs, skip the package entirely and use the client-side adapters — they move content over REST in either direction, including to the filesystem (see Fidelity).
Verification & reports¶
job.verify() reconciles item counts and spot-checks content checksums, and
job.export_reports(dir) writes Summary/Item/Failure reports as CSV + JSON.
See also¶
- Data pipeline — tabular import/export
- Large lists — thresholds that shape a migration plan
- Limits — the service limits the toolkit respects