Save Customer-Provided Data and Derived Outputs

When Workbench notebooks write CSV or similar file outputs with Spark, the output directory usually contains part files instead of one CSV file. Store these outputs in an approved destination according to how the output will be reused.

To prepared and save data for reuse:

  1. Before saving, confirm that you have approval to use the customer-provided data or to create the derived output.
  2. Confirm your approved destination for your workflow. Where configured, use approved Object Storage locations for customer-provided files, external tables, external volumes, curated files, or exported outputs.
  3. If you are uploading customer-provided data to Object Storage, sign in to the OCI Console and upload files only to the approved bucket, prefix, or quarantine location provided by your administrator.
  4. If an approved workflow requires copying a file from a user Object Storage location to a shared Object Storage location, use the OCI Console Object Storage copy action for the source object unless your administrator provides a different approved process.
  5. Confirm that the Object Storage location is connected to Workbench through an approved external table, external volume, external catalog, or workspace configuration.
  6. Select Save to save outputs to the approved destination according to your workflow.
Use the following patterns to decide where data belongs and where data can be discovered:
  • Autonomous Data Warehouse (ADW) Real-World Data (RWD) and Reference Data: ADW RWD is read-only data and is used as source data for authorized exploration and analysis. This data is available through the Master Catalog or in authorized read-only schemas.
  • Workspace Files: Workspace files are used for interim files, notebook inputs, small scratch outputs, or workspace-only analysis files. This data is available in your workspace, but is not automatically added to the Master Catalog or available in a user's Object Storage files.
  • Object Storage Files: Approved Object Storage locations are used for customer-provided files, curated files, exported files, or notebook-derived file outputs. This data is available in Notebooks with OCI configured access.
  • External Volumes: External volumes are used when files remain in an approved external Object Storage location and notebook users need file-oriented access through a catalog volume path. This data is available through the Master Catalog or through volume paths.
  • External Tables: External tables are used when approved files in Object Storage are queried as tabular data through the catalog. This data is available through the Master Catalog or can be queried with SQL or Spark table references.
  • Managed Volumes: Managed volumes are used when Workbench manages the file storage for a cataloged volume. This data is available through the Master Catalog but may not be available outside of Workbench.
  • Managed Tables: Managed tables are used when Workbench manages storage for tabular data that users query through the catalog. This data is available through the Master Catalog or can be queried with Notebook code.
  • ADW Tables or Schemas: ADW tables and schemas are used for configured tabular outputs, direct SQL workflows, approved reporting schemas, or Oracle Cloud database-backed datasets.

Note:

All data patterns are dependent on the user's access and configurations. External table and volume locations must be in a separate approved Object Storage prefixes.