Implement a scalable machine learning workflow using Oracle AI Data Platform and OCI Data Science
The solution combines the strengths of both platforms:
- AI Data Platform: Provides centralized and governed data management, large-scale Spark processing, workflow orchestration, and a medallion data architecture.
- Data Science: Provides a flexible runtime environment and infrastructure for notebook- and job-based machine learning development, model deployment, and model serving for production workloads.
This architecture is designed to support centralized enterprise data platforms, large-scale data processing, medium to large-scale machine learning workloads, and collaboration between data engineers and data scientists.
Typical use cases include:
- Retail forecasting: Generating daily demand forecasts based on both historical and newly ingested data.
- Insurance renewal optimization: Optimizing insurance renewal prices to maximize margins while satisfying business constraints. A new batch of renewal offers is generated each week.
- Telecom log event monitoring and drift detection: Analyzing large-scale operational event streams to detect data drift, identify anomalies, and generate daily reports and insights.
Architecture
This diagram illustrate this reference architecture.

Description of the illustration aidp-ml-workloads-architecture.png
aidp-ml-workloads-architecture-oracle.zip#GUID-73867B97-5BE4-434A-85C4-097DB840B323
The architecture is divided into four primary domains.
AI Data Platform (AIDP)
AI Data Platform is responsible for data ingestion, transformation, governance, and preparation.
- Data Integration: Ingesting data from diverse enterprise sources using various tools such as OCI Data Integration and OCI Streaming with Apache Kafka
- AI Lakehouse: Scalable and dynamic storage
- Spark Workflows: Scalable data processing and transformation using the medallion architecture with Bronze, Silver, and Gold layers
- Agent Flow: Natural language data interactions
These resources are primarily created and managed by data engineers.
Data Science Platform
Data Science manages the machine learning lifecycle, with different resources used during the R&D and production phases.
The R&D phase includes:
- Notebook Sessions: ML experimentation and model training
- Model Catalog: Model versioning and model management
- Model Deployment: Deployment for sclalable model serving
The production phase includes Scheduled Job Automation for repeatable batch inference pipelines and production automation
OCI Monitoring
OCI Monitoring provides centralized operational visibility across both platforms for monitoring, visualization, and alerting related to data drift and model drift. OCI Monitoring helps ensure operational reliability and production readiness.
Platform Interaction
The interaction between the Data Science platform and the AI Data Platform primarily serves the following purposes:
- The Data Science platform retrieves data managed by the AI Data Platform for both model training (historical data) and production automation (new cases requiring scoring).
- The Data Science platform enriches the data by adding prediction scores for newly processed cases.
The data layer used for training and production automation may vary across applications, depending on the schema and architecture defined by the data engineer. In many cases, the data is sourced from the curated Silver layer, but it may also come from the Bronze or Gold layers, or from a combination of multiple layers.
Similarly, the enriched data containing prediction scores is often written to the Gold layer, but it may also be stored in other data layers, depending on the application's requirements and data architecture.
The data used for data drift detection can originate from the Bronze, Silver, or Gold layers, depending on the specific use case and data architecture.
Architecture Components
- Oracle AI Data Platform
Oracle AI Data Platform is a unified platform that simplifies the cataloging, preparation, and analysis of data across your data estate. It brings together data, AI, analytics, and governance within a cohesive user experience enabling you to build secure, scalable AI-powered applications. Oracle AI Data Platform unifies Autonomous AI Lakehouse, Oracle Analytics Cloud, OCI Object Storage, OCI Generative AI and Fusion Data Intelligence.
Within this platform, Oracle AI Data Platform Workbench provides a dedicated development environment for you to design, orchestrate, and deploy data pipelines and models, set RBAC policies, and use open source technologies such as Spark to prepare, analyze, and enrich your data.
- OCI Data Integration
Oracle Cloud Infrastructure Data Integration is a fully-managed, serverless service for designing and executing data pipelines. It enables seamless extraction, transformation, and loading of data into OCI targets like Autonomous AI Lakehouse and OCI Object Storage. Users can build integration flows through a codeless, intuitive interface that auto-scales execution environments. It supports both ETL with Spark-based processing and ELT using SQL Pushdown for performance and efficiency. The service also offers tools for data preparation and protects against schema drift with rule-based handling.
-
Data Science
Oracle Cloud Infrastructure Data Science is a fully-managed, serverless platform that data science teams can use to build, train, and manage machine learning (ML) models on OCI. It can easily integrate with other OCI services such as Oracle Autonomous AI Lakehouse, Oracle Cloud Infrastructure Object Storage, and more. You can build and evaluate high-quality machine learning models that increase business flexibility by putting enterprise-trusted data to work quickly, and you can support data-driven business objectives with easier deployment of ML models.
- OCI Identity and Access Management
Oracle Cloud Infrastructure Identity and Access Management (IAM) provides user access control for OCI and Oracle Cloud Applications. The IAM API and the user interface enable you to manage identity domains and the resources within them. Each OCI IAM identity domain represents a standalone identity and access management solution or a different user population.
- OCI Monitoring
Oracle Cloud Infrastructure Monitoring actively and passively monitors your cloud resources, and uses alarms to notify you when metrics meet specified triggers.
- OCI Object Storage
OCI Object Storage provides access to large amounts of structured and unstructured data of any content type, including database backups, analytic data, and rich content such as images and videos. You can safely and securely store data directly from applications or from within the cloud platform. You can scale storage without experiencing any degradation in performance or service reliability.
Use standard storage for "hot" storage that you need to access quickly, immediately, and frequently. Use archive storage for "cold" storage that you retain for long periods of time and seldom or rarely access.
- OCI region
An OCI region is a localized geographic area that contains one or more data centers, hosting availability domains. Regions are independent of other regions, and vast distances can separate them (across countries or even continents).
- Service
gateway
A service gateway provides access from a VCN to other services, such as Oracle Cloud Infrastructure Object Storage. The traffic from the VCN to the Oracle service travels over the Oracle network fabric and does not traverse the internet.
- OCI Streaming
Oracle Cloud Infrastructure Streaming provides a fully-managed, scalable, and durable storage solution for ingesting continuous, high-volume streams of data that you can access and process in real time. You can use OCI Streaming for ingesting high-volume data, such as application logs, operational telemetry, web click-stream data; or for other use cases where data is produced and processed continually and sequentially in a publish-subscribe messaging model.
- OCI virtual cloud
network and subnet
A virtual cloud network (VCN) is a customizable, software-defined network that you set up in an OCI region. Like traditional data center networks, VCNs give you control over your network environment. A VCN can have multiple non-overlapping classless inter-domain routing (CIDR) blocks that you can change after you create the VCN. You can segment a VCN into subnets, which can be scoped to a region or to an availability domain. Each subnet consists of a contiguous range of addresses that don't overlap with the other subnets in the VCN. You can change the size of a subnet after creation. A subnet can be public or private.
Recommendations
Your requirements might differ.
- Automate repeatable inference pipelines by using Data Science Jobs and schedules.
- Use Data Science Deployment and Serving tools to support scaling, operational monitoring, and reproducibility.
- Monitor production machine learning workloads for operational health, data drift, and model drift.
- Establish clear operational ownership between Data Engineering and Data Science teams.
Considerations
- Implement OCI Identity and Access Management and role-based access control (RBAC) policies across both platforms.
- Monitor production workloads for operational health, data drift, and model drift.
- Use scalable governed storage for long-term data retention and analytics.