WEBINAR Series

Optimizing Your Data Strategy with Modern ETL

Use Salesforce regularly? This webinar recap is for you. Here, Integrate.io's panel of experts explore hot-button Salesforce issues and more.

Optimizing Your Data Strategy with Modern ETL

About the speakers

Mark Smallcombe, CTO, Integrate.io

Mark has more than 20 years of Technology and Software Engineering Executive-level Leadership experience including CTO, VPE, and Director-level positions. I have a track record of success in assembling, leading and working with engineers to build consumer-based products that compete successfully in the marketplace.

Noel Yuhanna, VP Principal Analyst, Forrester

Noel has over 25 years of experience in IT and has held various technical and management positions. He has taught several technical and management workshops on big data, data management, data integration, building scalable apps, and more!

VIEW TRANSCRIPT

Welcome everyone. Thank you for joining webinar on optimizing your data strategy with modern ETL.

First, a quick introduction of today's speakers. My name is Mark Smorkum. I'm CTO of Integrate.io. I've led product and engineering teams in Silicon Valley, Los Angeles, and Sydney.

Our special guest today is Noel Johanna, who's VP and Principal Analyst at Forrester.

Noel has a wealth of knowledge of big data and data integration and has written many Forrester reports on these topics.

Here is today's agenda. I'll give a short introduction to X Plenty.

Noel will then talk on leveraging data pipelines to accelerate analytics.

I'll then talk on a promising new architecture, ETLT, and show a demo of Integrate.io.

We'll then move into Q and A for ten minutes to close the session.

We will be making this recording and the slides available after the webinar.

What is Integrate.io? Integrate.io is a cloud based low code ETL and ELT platform that has a really nice drag and drop user interface for creating your data pipelines.

We also have an API that can be called from Apache Airflow.

We support all common data sources and data destinations and have over one hundred pre built data transformations for you to use.

We lead with our fantastic support team. We have a global support team of experienced data engineers ready to help build your pipelines.

And we have many happy customers from small startups right up to big enterprises.

I'll now hand over to Noel Yohanna. Noel, please take us into the first part of this session.

Thanks a lot.

Thanks, Mark, and great to be on this webinar and talk about data pipelines. This is one of the, I think, hottest topics we have come across, especially as organizations start to leverage the cloud platform, right? So we're seeing data pipelinings becoming more important, especially for doing this modern day analytics, but also accelerating those analytics as well, right? And people also call this a modern ETL, and we'll get into discussion about ETL later as well, but data pipeline is definitely something which we hear from a lot of customers. I speak to about three to four customers every day from financial services, retailers, healthcare, government agencies, manufacturing companies across the globe. And people are dealing with all those migrations going on in the cloud, I guess, right? And so I'll share some of the things as to what we are learning from this, you know.

Now, what's interesting is that data has become the most critical asset for any business today to succeed in today's digital transformation.

You know, without data, the companies may have so difficulty in existing today, I guess, in this modern digital transformation era, right? With data, what can you do? Well, you can do a lot of things, right? You can increase your revenue actually in every industry, whether it's, you know, financial services sector, retailers, healthcare.

It's also improving customer experience. I think customer experience is a very important topic, especially in the retail sector, but also other sectors as well. You want to be making sure that people are getting the value out of that, the product and services you're selling, right? How can we understand the customer, right?

And it's all about data, right?

Retaining customers is critical as well, right? You want to retain customers as well. A lot of customers actually have choices today. They can go to the internet, browse here and there, and say, by the way, I'm going use this services and that products, because they are better, right, and hence you may lose a customer, right, churn happens as a result of that, right.

It enables innovation, I think innovation I think is very important. Today's day, innovation cannot happen without data, hence the data driven enterprises have emerged, right? We have seen those who have been focusing on this whole data driven are driving more revenues, better customer success rate, better, obviously, competition in terms of being able to provide better value to these customers. So innovation definitely is a critical component, and, you know, data provides you that context, new products and services, expanding markets, and many more, right?

So, I mean, without data today, I mean, organizations would not succeed. In fact, I think every organization today or every company today is a data company, actually, besides whatever they sell and what have you, right? Fifty one percent are making money with data, which is very interesting. You're buying and selling data as an asset, right?

So that's why it's very important, it's become the most important asset today, which can determine what's going on, right? People want to buy and sell data to understanding their customers as well.

There's twelve zettabytes of data that exist today, twelve zettabytes, that's about twenty one zeros. I don't think it'll fit into my iPhone actually, if I try to put that in, it won't fit in, right? So that's twenty one zeros, by the way, so there's a lot of data there. We got data on premises, we got data in sensor devices, we got data in cloud, edges, and also SaaS applications as well. Now, more data is good, by the way, right, because the fact that more data means new business opportunities for every organization.

Now, people were dealing with CRM and ERP systems for decades, but today it's different types of data, as I mentioned, right? Could be social data, your Twitter data, your Facebook data, your LinkedIn data.

You want to leverage that data as well for doing customer intelligence, minimizing customer churn analysis, right, to doing churn analysis.

So a lot of the data is existing there, so organizations should definitely be leveraging the data to driving insights.

More organizations are moving towards, you know, cloud to lower the cost and becoming more agile. You know, over the last, well, this year, due to COVID-nineteen, we've been getting inquiries from customers who are trying to move to the cloud. One of the big reasons is the cost, right? You can save money.

A lot of organizations save more than twenty percent or more when you move from on premises to the cloud, you know, to do, supporting these different types of insights, whether they're doing customer intelligence, customer three sixty, or IoT, and what have you, right? So that's a good thing.

Why does the cloud bell? It's not just about the cost, it's also agile, it's automation, it's security, it's scale, it's serverless, you know, the list goes on and on, right? Can you scale your platform for sixteen CPUs to thirty two CPUs in a second?

Yes, you can do that in the cloud, not on premises. You probably get you six months to get you the hardware hardware shipped to your destination, right? So that's the big problem. You know, cloud offers you agility, right, and automation. All the software automation that's happening today is in the cloud.

That also is a very important point to consider as well.

You know, so the modern day platforms, you know, as we call the cloud platforms, have emerged, and cloud platforms are enabling to supporting this new and emerging insights quickly. Now, is a modern day data platform? This modern data platform is all about providing you that level of ingestion, transformation, integration, and also storage, as well as the consumption could be, which is through visualization and dashboarding as well. So if you look at this diagram, there's data pipelining, which is a very important component of this modern day platform, which includes the ingestion of SaaS data, CRM data, social data, clickstream data, and all these other types of data sources coming into, in real time, near real time, and streaming sets as it gets into the data warehouse, which could be in the cloud, and it gets transformed as well.

So data gets transformed and then data gets into you're converting data into insights at the data warehouse layers through the transformation process, but data warehouses also are running on this elastic scale model, right? You can scale, you can build a data warehouse with multiple petabytes in minutes, in seconds, that's how fast a cloud data warehouse can be built today, right? Automation and AI machine learning also is there to automate the processes of administration, backup and recovery, disaster recovery, you're doing security, all these could be automated within the warehouse platform as well.

And then you've got the actionable insights, which creates all those dashboards, the visualization, as also creating possible actions as well. So this whole end to end has become critical, and you know, the good thing with this modern data platform, especially the cloud, it's agile, as I mentioned, right? So if you look at all these benefits you get, you get the automation benefits, you get the scale benefits, you get the security benefits, you also definitely get real time benefits as well, which I think is a very important element as well.

And then it's got integrated multi persona, time to value, and lower cost. Cost is obviously a factor because you want to get it end to end, but also it's integrated, which is very important. A lot of companies are spending time in getting data and transforming the data into warehouses with scripts, with programs, they're writing C programs, they're writing SQL statements, and sometimes they run, sometimes they don't run, sometimes they don't scale very well, and that's a big issue, right? And data pipelines is all about that scale, which you can ingest this data quickly enough, transform this data, load this data into the warehouse, and put this more actionable as well.

And multi persona is very important. You know, previously, most of the IT platforms were only managed by, well, managed by IT and used by IT, except for applications, right? But today it's different, Data platforms is being used by business analysts, business users, business community as well, which is a changing kind of a transformation that's happening in the industry, which I think is very important, because you need all these platforms to be self-service driven, which I think is going to be critical as well, right?

And as modern day platform use cases continue to grow, we got Customer sixty, churn analysis, upsell, cross sell, ad campaigns, promotion planning, new launches, and what have you, right? IoT analytics is what people are doing with these modern day platforms.

Detection, risk analysis is also being done actionable real time insights. I think actionable real time insights is very, very important.

A lot of data we have moved around has been batch oriented.

Hence, hence we are getting reports twenty four hours late, right? Why? Because because it's batch oriented stuff going on, which is slow, slowing down these insights, and today, imagine your email is coming twenty four hours late, will that be acceptable? No, right?

That's exactly what the business issue is today. You want to be able to drive these insights in real time as well. Line of businesses focused analytics like sales, marketing, finance is also something which we are seeing customers are demanding more as well. And data platforms are delivering that value proposition.

It's able to do the data collaboration on a platform so that even business analysts, business users, could actually use data for dashboarding and visualization as well.

So if you look at it in terms of data pipelining in modern day ETL, right, and I'll talk a bit more about ETL, but the data pipelining is all about ingesting, processing, transforming, high velocity of data for integration flows, data movement, persistence, in a self-service, very important, self-service and automated manner, right? We can move data around, we have been moving data around for decades, right, for forty years, but the question comes back is, can it be doing this whole movement in a more self-service automated manner, in a high velocity, right? That's the biggest challenge today, right?

And plus also being able to transform this data, being able to process this data in a very highly parallelized environment for high performance, which is what this data pipelining is all about, right? So data pipelining, if you have not invested in data pipelining, you should be doing it today, because it's a very important investment, especially as you go to the cloud, especially as you go to the multi cloud, right? People are not having just one cloud. They got AWS, you got Azure, you got Google, you got also a hybrid cloud, you got on premises as well, which is the things running of data as well.

So you need to actually have a better facility as well, right? So, you know, ETL has been around for more than, you know, forty years. In fact, I think I think the first reference of ETL came in 1970s.

I personally use ETL back in the late 80s.

And, you know, data movement was there in the warehouses when we were trying to build these databases in the late '80s, and it's interesting, like, obviously ETL has emerged a lot over the last many years, and even though now we talk about data pipelining, and the real reason is because it's more modern day architecture, it's able to build on terms of Kubernetes and containers, it's able to do real time, it's able to do all this automation, which really the difference is then what ETL traditionally have done from a technology perspective. So if you look at concepts, right, ETL concepts have been there for how many years? Well, right, many, thirty, forty years, and ETL has been there, the concepts are still there, even for data pipelining, the concepts still exist, right?

And the concepts for ETL in terms of data pipelining is when you're extracting data from these source systems like databases, like applications, like warehouses, and you transform the data, you wanna transfer it for security reasons, you may wanna transfer it for data consistency reasons, you may wanna transform the data for making sure it is uniform in terms of its usage across data sets, right? Because sometimes when you load data, it's like garbage in garbage out actually.

The pipelines have to be streamlined, know, sometimes you don't want to dump data into a warehouse, which is not great, you know, right?

So that's the reason why transformation helps to get consistency, to get also security control as well. You may have PII, PHI data moving on those ETL frameworks of yours, and you're wanting to load this data as well, right? The load part of it is going to the warehouse most of the time, right? So that's what ETL does, extracting the data from the source system, transforming the data, and then loading it into a warehouse. ELT has been around as well, right?

ELT has now gained more traction over the last decade as big data came into existence, right? You wanted to load the data into the lake, data lake, quickly enough, and then transform the data, because lakes are very fast in transforming data themselves, right? You can run those processes in the lakes as well, plus the data is available in more real time nature to extract the data and you push it out into these data lakes. You can also do it for data warehouses, destination could be lakes, it could be also warehouses as well.

And then we have the best of the breed, best of the both world, ETL, ELT together, we call it ETLT, right? The industry has been using this terminology for a while, and this is extracting, transforming, loading it, and then transforming, and then why do we care about this? Well, you care about it because you may be loading data, is sensitive in the data warehouses, PII, your credit card, your social security numbers could be getting into the warehouse. How do you stop that from doing it?

Well, you can transform the data, you can de identify the data, you can mask the data, like credit card data, before it gets loaded into the data warehouse environments. So this becomes very important as well. Plus also, you can transform the data so that you can eliminate data, which is not wanted, You can eliminate the data. And this is a very important point for organization as well, so you not want to You wanna throw away data sometimes because the data may be just not the right data you want to load.

That's what transformation can do with this as well. So these concepts are very still applicable in this data pipelining scenario, so when you're building the data pipelining, which is the collection and processing and transforming the data, you can still apply those concepts into this framework.

And you know what data pipelining does? As mentioned, real time automation, self-service, security, agility, all of these are important points. Sources could be any sources, right? Database, warehouse, legacy, social data, and the destinations could be dashboarding, warehouse, exploration lakes and what I mean, without this, it's a very difficult world altogether.

Right? In fact, if you look at data engineers, right, what are data engineers doing? They're doing data pipelines. In fact, a lot of organisations are spending a lot of money today in building data engineering teams.

And what are the data engineering teams doing is they're building these pipelines for all the use cases I discussed customer three sixty, customer intelligence, IoT analytics, fraud detection, and what have you, right? So it's kind of interesting kind of a thing of how this comes together.

So, you know, people always ask, okay, so how do you compare traditional ETL, right, which has been around for thirty, forty years, and how does it stack up against the data pipelining, which is, some people also call it more than ETL, but data pipelining is a terminology being used, more acceptable terminology, because it does differently than slightly, more than ETL actually, right? ETL is there, but it does more than ETL, and if you look at this diagram, speed has been mostly batch for ETL, it's more real time streaming, but also can do batch if you wanna do batch, but agile being not so agile versus pipeline, we're more agile, self-service sources of data has been mostly traditional database stuff, all kinds of data, and cloud to limited sources, and you you get the idea, and there's scalability, scalability has been good.

I mean, I don't think ETL is slow, it's just that real time is better in pipelining, right? So multiple personas, ETL requires a lot of coding, right? You've got to code, you've to write programs, sometimes you've policies. Here, you can just drag and drop, drag and drop.

Even a person who does not understand the concepts or even doesn't understand programming can do data pipelining. That's a big difference. It can do ETL, ELT, security and ease of use, I think is better better as well, right? So a lot of benefits, which I think is really the factor as well.

You know, I would say data pipelines have become critical, right, for cloud, especially cloud, right? Why it builds this new generation of analytics insights quickly, such as Customer three sixty, Realtime Insights, IoT. So especially as you're building this new generation of analytics, you want to be able to drive these types of data pipelining scenarios, automate the ingestion and processing workflows through zero code. This is very important zero code low code requires minimum learning, and I think I think you know no one has got like six months trying to build these pipelines anymore.

In fact, ETL developers used to be there right many ETL developers, and you know, when you ask them, okay, I want to move data from A to B, yeah, come back next week, we'll give you the programs, make it ready in a week's time. What if I want run this in the next minute, right? You don't have time to write programs anymore. That's where data pipelining comes in, right?

It automates that whole process with zero code, low code, right? You can drag and drop those frameworks. This is exactly why this is changing the way we do data processing, and especially the automation that comes in, and we're also starting to see a lot of machine learning and AI intelligence being built into the pipelines so that you can actually help become more intelligent scenarios as well. The platform becomes more intelligent to understand the data coming in and processing as well.

Helping supporting various personas, whether it's developers, it's business analysts, data engineers, power users that want to move, process, integrate data, right? It's not just about, you know, ETL traditionally has been more on the ETL developer side, administrator side, but now we are also seeing data pipelining expanding the scope to reaching various other developers, but also analysts, business analysts. We have never ever in the history of data movement have business analysts moving data, especially at the enterprise scale level, and that's the big change in this industry. Data engineers and power users are using this as well, they want to be able to move and process and integrate data as well themselves, and this is where we are heading towards.

We are moving from IT, information technology, which we have had for decades now, towards business technology.

Business technology is all about empowering the business users to drive better success themselves, instead of relying with IT organizations to help them in building some of these things, right? So that's the big change, what we see. Data pipeline, you can deliver data to various data stores, including databases, warehouses, analytics. Mean, the good thing is it's common framework to doing many different things, right?

Even data collaboration could be enabled through pipelines. And they leverage new data sources easily to support I think this is a very important point, being able to leverage new data sources to driving better analytics for various industrial analytical scenario as well. This is, I think, a very fundamentally important area where we see customers investing in because the new sources are coming in, you got edge data coming in, you got all the sensor data coming up, you got also these device data, log stream data, you got the other SAS data, you got Twitter and Facebook and LinkedIn, you got all these other types of data coming in as well, Big data, small data, you name it, right?

And so you need actually a platform that actually is more adjustable, more adaptable towards these new generation of things, and that's what data pipeline does in deliver value.

So what are the recommendations I have? Well, data pipelinings should be part of your data and analytical strategy, if it's not already, right? I mean, so especially as you go to the cloud, start to look at data pipelining, right? You know, people are building all these various data and analytics strategy, definitely look into that data pipelines, right?

And, you know, this is an important area, as I mentioned, especially as you build an agile, trusted, self-service driven platforms to doing all these use cases which people are building, right? Data pipelinings helps automating the ingestion, processing, the transformation when moving data from sources to data warehouses and data lakes. You know, they are able to help automating, right? This is the biggest thing, the automation piece, right?

Not that you're writing code, you're not that you're writing things which don't break, right, as much. But the more important thing is, is that whole automation piece, I think is really driving success. Not only customers are able to do this quickly, but they are saving money, right? Saving money on resources, saving money on additional compute and resources that may be spent.

I know one customer was mentioning to me that they were they were doing some traditional way of data movement. And you know, they weren't having a staging server, another server to move data before it got caught into a warehouse. And they had like two, three different servers, and it was just slowing slowing them down, but also adding to the cost of the compute and storage, besides slowing down everything, right? So data pipelining is really that process which really helps you in automation.

Data pipelining can, used by multiple personas, including data engineers, analysts, as we talked about it, even security people. The good thing with data pipelining is that you have a variety of personas that can help you, can leverage the platform. Data pipelinings can ensure data quality, data security to protect and deliver trusted data. You know, data quality is a very important discussion, especially, you know, as you move data, as I mentioned, the transformation, the ETL transformation part is where data quality can fit in as well.

People build a lot of these policies, but also change the data. Sometimes you only have, you know, addresses may not be completely called out like a street, maybe called an ST instead of street, right? You may want street completely called out in the data warehouse, hence the transformation can do that as well, right? Data quality checking to ensure that data is protected for GDPR regulations or CCPA, California protection as well.

So, I mean, you know, there's all these regulations that you have to ensure protection as well. Data pipelines can ensure that unauthorized people do not see the data like credit card data and source code numbers, so PII, PHR data could be protected as well before it gets into a warehouse, because a warehouse may have more users trying to get to the data, right? Look at data pipeline solutions. Solutions could definitely help you in improving productivity through automation and accelerating the use case as well, right?

So solutions can help. As I mentioned, you can manually do this, right, writing scripts and SQL statements and all this shell scripts and all those things, which can break, and people complain about that all the time because it breaks in the middle of the night, trying to fix it.

Data piping solutions are the ones which really automate that, but also can improve productivity level, right? You don't need so many different so many data pipeline engineers because of the need for simplification, which I think is a really value proposition going forward.

With that, I would like to thank you and hand it over back to Mark.

Great. Thank you, Noel. That was really insightful and informative.

One thing that stood out for me was the enormous amount of data now being generated by organizations.

I found the description of modern data pipelines very memorable, and everyone can benefit from your thoughts on modern ETL.

Now, let me share my topic on ETLT, Protecting Sensitive Data.

Consider this, is ETL dead?

Is ETL still relevant with ELT and data lakes?

Think about that for a moment.

Now, this leads me to the bottom line. Number one, ETL is still very relevant to clean and secure your data.

Number two, ETLT gives you the best of both worlds. Gives you ELT speed and flexibility and protection of your sensitive data.

Let me repeat that bottom line.

Number one, ETL is still very relevant to clean and secure your data.

ETLT gives you the best of both worlds. ELT speed and flexibility and protection of your sensitive data.

Today, let's review the pros and cons of ETL, ELT and discuss a new emerging architecture, ETLT.

In an ETL architecture, you extract data from databases, files, SaaS applications, and then you can transform this data using a platform like Plenty, and then load this prepared data into your data warehouse.

From there, you can generate reports and dashboards for your analytics.

It's a very popular architecture, and it's been around for many years, decades.

Transformations can protect sensitive data like names and email addresses and other personal data, and you can transform this data at its source. As an example, you could do this transformation in Europe for GDPR, and you only load what's needed.

The downside is it can be batch and but nowadays, batch is pretty fast. It can be a one minute to five minute delay depending on the transformations. So it's fast batch.

In an ELT architecture, you extract and load data directly to your warehouse.

It's really a copy of this data.

And then once it's in the warehouse, you can perform SQL based transformations on that data.

It started becoming popular a few years ago as cloud warehouses started becoming more performant and lower cost.

It's very fast at ingesting raw and unstructured data.

And the warehouse becomes more like a data lake.

More people can write these SQL transformations as well, rather than just data engineers. So it opens up the transformations to the whole team.

The downside is the data is not cleaned before loading.

And it can be bad if you have high volumes of data because your storage costs can increase, and then you have to put data retention policies in place.

And that can lead to a loss of this historical data in the warehouse, which is one of it, one of its advantages.

But the big disadvantage in this architecture is compliance.

ELT often copies sensitive data into your warehouse, and this can lead to accidental data breaches.

You might find sensitive data in your warehouse logs.

Someone could create a report that accidentally exposes sensitive data, or someone could download sensitive data to a laptop and then potentially lose that laptop.

And there's a growing list of compliance needs. GDPR, HIPAA, CCPA. It's an endless list of data protection privacy acronyms, and there's new ones coming out every day.

And there are scary fines for data breaches. Uber, one hundred and forty eight million.

Equifax, five seventy five million.

Yahoo, eighty five million, all for data breaches.

It's critical to protect your sensitive data before it gets into your data warehouse.

So what is the solution?

ETLT.

It's a blend of ETL and ELT, And it has two transformation steps.

There's a light transformation step at the beginning, the data protection, where you can remove and encrypt sensitive data. And you do it in a platform like Integrate.io. And then this cleansed but still raw data can be loaded into your data warehouse where you can do SQL based transformations in the warehouse.

This gives you much better data protection and significantly reduces your business risk.

You can remove and encrypt sensitive data before it leaves Europe for GDPR compliance.

ETLT is ELT with data protection.

I'll now give a demo of Integrate.io.

In this example, my business analyst wants some European sales data available for centralized US reporting, but we can't copy it directly because of GDPR, so we need to take some steps to protect the data.

This is X Plenty.

On the left hand side, have the navigation. We have connections.

You have packages where you create your data pipelines.

You have a way of monitoring your jobs that are running. You can create your clusters for those jobs to run on.

You can create schedules.

And then in settings, you can invite your teammates into your account.

So we'll just look at new connection and just see the type of connections we support. Lots of analytical databases, file stores, NoSQL databases, object stores, relational databases, many services.

One of the popular connections we have is this REST API connection, which allows you to connect to all sorts of systems and is often used to connect to ERP systems. And the Salesforce connector is really useful as well as it's bidirectional.

So we're going to go into packages.

And we'll create a new package and we'll give it a useful name.

So let's say European leads to the US.

We'll add a component.

Here, we're gonna add a database, and I already have a connection set up.

So that's my MySQL database.

I have a table in there called leads.

And now it's automatically pulled in the schema, and it shows you all the the data or a preview of the data that's in there. And as you can see, we have some sensitive data. So we have names of the leads. We have the email addresses. And we also have the IP address as well. So we want to start removing or encrypting some of this data.

So we'll select all the fields that we want.

That looks great.

Now we'll add a component. These are table based transformations. So we'll add a select component here.

We'll auto fill it with all the fields that we had before.

Now I'll put some transformations in here and they're a little bit complicated. So I'll just copy them first and then we'll go through them one by one.

Okay. So if we look at the first one here, what we're actually doing is encrypting the name field. And we're using the customer's Amazon KMS service to do the encryption. So we ask the KMS service for an encryption key.

We then encrypt it with envelope encryption. And we actually use AES two fifty six encryption. So it's it's very secure, and there's a documentation on it. So that looks great.

On the email, what we want to do is remove the first part of the email address. So we're doing a string split on that, and we're splitting on the ampersand, and we'll just leave the domain on of the email.

And then for the IP address, what we're gonna do is remove the last octet of the IP address. So we'll keep the first three, but remove the last one.

And, that makes it impossible to actually track it down to an individual user.

So that all looks great.

Now we'll add in the warehouse that we want to put this into. So we're gonna use Google BigQuery.

Connect that up.

And we'll put it in the same table name, but in BigQuery. And we're gonna actually override the data that's in there.

That looks good.

And we'll autofill it with all the fields that we want.

That's And then fingers crossed, it saves and validates.

Great. No errors. That's perfect.

So now what we'll do is we'll run that job.

And I have a cluster here sitting in Europe where we can actually run that job.

There's some settings we can go through, but we'll we'll just run it.

Once that's running, I'll show you how you can schedule your jobs as well.

So if you go to schedule, we can create a new schedule here.

Again, we can give it a a useful name.

We'll give it the same as package. We've got lots of options to configure when you wanna run this schedule. You can run it every minute if you wanted to. We run on hours, days, weeks, months. You can also select whether you want to use cron or this drag and drop, this sort of interface here.

You can configure how big you want your cluster to be to run this job and how long it takes for it to terminate that cluster.

And then we can add the package we want to run.

So we'll add our package.

And we're always going to run the latest version of the package.

Some teams, if you're constantly developing your package, we have versioning. So you can actually run a baseline version of your package if you want to, so you can continue developing without impacting any schedules that are running. But we'll keep that.

So that's great. So the schedule is now created.

And what we'll do is we'll go back to our jobs.

And yes, that job has now finished.

We can quickly look at the details. Yeah. It's written twenty two records to Google BigQuery.

And if we flip over to Google BigQuery, you can actually see the records that have come in.

So we've got the name, which is now an encrypted blob of ciphertext.

And we have the company name. We have an email. It's now being converted to just the domain name.

And the IP address has got the first three octets and not the last octet. So that looks pretty good, but it's a little cryptic for our for our business analyst. It's not not super friendly. So what we'll do now is we'll just do the last transformation and we'll do that in SQL.

We'll use BigQuery to do that transformation. And here what I'm doing is actually just hashing the the encrypted ciphertext to create a nice user ID for the business analyst. And I'm changing the email over to to a domain. And I'm also mentioning that we've actually truncated the IP address.

So if we run that, it should clean up the data a bit. Oh yeah, that looks a lot better. So now we've got a nicer friendly user ID for the team. We've got the domain here and we've got this truncated IP address.

And if I wanted to, I could save that as a view and then it would it would run automatically.

So that is great. Now I'll flip back to our slides.

So that concludes the demo. And what I've shown you is really the full end to end ETLT experience. So we started off with the ETL and then we did the last transformation in BigQuery for the last t of the ETLT and that cleaned it up for the business analyst.

Now let me go back to my initial question.

Is ETL dead?

Let's agree. ETL is very much alive and critical for data compliance. So let me recap the bottom line.

ETL is still very relevant to clean, process, and secure your data. And ETLT gives you the best of both worlds. It gives you the ELT speed and flexibility and protection of your sensitive data.

Thank you.

We'll now go to q and a.

So I have a few questions that have come in already.

So Noel, how do you see modern data pipelines integrating with machine learning technology?

And are many companies doing this already?

Yeah, well, I mean, there's obviously machine learning technologies which are still evolving. When it comes to AI machine learning, we are very early on in this model, because the automation so so let's put that into perspective. AI machine learning is two sides of a coin.

The first side of the coin is all about the consumption side, where you actually have business users, business analysts, businesses leveraging data, and putting machine learning models to be able to understand data for their customers, for their businesses, and what have you, right? So that's one part of that whole dimension of AI machine learning, which has been there for a couple of years now, and machine learning models could be built for that as well, right? And people have tools to help you build that scenario. The second side of the coin is the machine learning AI built into the platform itself, like a data pipelining, to be able to understand what data is being ingested.

Imagine data pipelining, you actually get your data into it, is this credit card data coming in automatically, instead of someone looking at the screens and spending countless hours and maybe years trying to figure this out, you only have built in intelligence into the pipeline to understand what data is there, PII, PHI, credit card data, social security numbers coming in automatically, right? That is what machine learning and AI is all about in the pipelining scenarios. It's also helping you to determine what kind of data connections need to happen within these pipelines. You may be having forty pipelines coming in.

Some of these pipelines may be interrelated, some of them may not be, because on certain intelligence, you're be able to provide that level of integration as well, right? And this is what we see as a big element where there's a big driver towards AI machine learning, especially around this intelligence, or what we call adaptive intelligence, around this data management layer, which is ingestion, transformation, governance, security, preparation, coming together from end to end perspective, and this is where the industry is moving towards when it comes to these newer generation of use cases. So we are very early on this phase of AI machine learning, but definitely there's a lot of opportunities here, and organizations should definitely be leveraging it in terms of their scenarios. Mark, what do you see in your expertise solutions of how machine learning is helping customers?

It's still, in our side, it's still very early. So we do have some customers that are using machine learning, And they're using, they're actually using Apache Airflow to do sort of workflows. And then they're calling the ETL and then doing machine learning type workflows using Apache Airflow. But it's still very early. Think a lot of people aren't really using machine learning yet.

Absolutely. I mean, as I mentioned, you know, is very early on, I think we are only like fifteen twenty percent in terms of its capabilities of machine learning within a data management, data pipelining framework, we have to yet see a lot of other benefits coming in into these pipelines, and I think, so this is going to evolve, the market is going to evolve, but you know, as you get towards this more self-service data pipelines, right, this intelligent data pipelines, we're going to see a lot more of that intelligence coming in, which means that the system automatically detects what data is coming in, the system automatically connects the data, it also knows what data to be throwing away, instead of you actually trying to throw away data, the system will throw it away for you, actually, especially if you have billions of records. Humanly, it's not possible to view billions of records, which may take you one hundred years trying to figure this out.

But that's why this intelligence of AI machine learning is coming in to really drive some of those insights quicker.

Yeah, no, that's great. All right, next question I've got, Ken, has come in.

How do large multinationals deal with the huge data volumes and potential data latency? And how do they do they process the data in each country before centralizing into a warehouse?

Yeah. Mark, what is your opinion on that?

Well, we see it. Yeah, we definitely see people doing their transformations in Europe for GDPR, but we actually have other ones that do it for compliance in other regions. So we have a data center in Frankfurt. And also even in Australia, they need to be able to do the transformations in a country for compliance. That's And then, yeah, we find the multinationals often do do the transformations remotely, and then centralize it in a, you know, a BigQuery or Redshift or something like that in the US, because a lot of the multinationals seem to be US based, you know, our customer base anyway.

Yeah, I mean, what we see as well, you know, the whole notion of GDPR regulations have put a lot of pressure on organizations, especially data movement. And what's interesting is that, you know, there is this notion of dynamic data masking, which exists, and dynamic data masking is all about moving data and de synthesizing the data, de identifying the data in real time as you move data, as you access the data in real time. It's not persisting, it's not physically masking the data, but it's actually doing it on the query or the access point, which are really helpful for data masking because this dynamic data masking actually helps a lot of these companies, especially with GDPR regulations, so that, you know, when you move data around, you are masked the data, right, and then it moves around, which is perfectly fine, you know, and I think so that's one thing which we're seeing as part of the whole thing, and you can do that dynamic data masking as part of that whole data pipelining as well, some lot of companies doing data masking as part of data pipelining, so that's a transformation step actually, right?

When you extract data, you're transforming the data, the transformation is the time when you're going to be doing the masking of data, and then pushing it out into a warehouse or into a Tableau visualization, perfectly fine, you know, right? And that's what we see as a whole platform from end to end. So yeah, I mean, it's a big concern, you know, security and governance, and I think even data lakes, I think, is interesting, you know, we only see about fifteen to twenty people accessing a data lake in an organization, like why fifteen, twenty people only? Well, it's only IT organizations why because when you're dumping the data into a lake, no one has a clue what data is being dumped into the lake.

The problem is that it could be sensitive PII PHI data being dumped into the lake, and someone actually has to go and figure this out, actually. So there is a raw data lake, and there's a curated data lake, which gets moved from the raw into a curated data lake, and that's what we see the governance playing a big role in this environment. So, you know, people have struggled for data lakes, well, security in the lake, especially with not knowing what data gets into the lake, and I think data pipelining is definitely helping in meeting these requirements because you can actually de identify those data, especially with consumer data, which needs to be done, you know.

Okay, that's great.

Next question. There seems to be a lot of innovation in the warehouse space.

You know, obviously Snowflake's IPO and there's a lot of movement in that space. How do you see that impacting modern data pipelines? And will warehouses start vertically integrating to create an end to end data experience?

Yeah, I mean, you know, and I think data warehouses definitely have been expanding in towards end to end, right?

Earlier this year, we did a Forrester Wave on data management for analytics, and this data management for analytics is all about this end to end, right?

People don't really care about whether it's running on data warehouses, Kubernetes.

Like, when you put on your Netflix, you get a movie, right? Is but guess what? Netflix is running on Kubernetes. Do you do you really care about that? No.

What I need is my dashboards, my visualization, my revenue up, and that's what is happening, right, in the industry. It's all about this end to end thing which becomes faster time to value.

You know, so agility, time to value, which I talked about, is critical for organizations going forward. We are seeing a lot of these companies focusing on that faster time to value, and so end to end becomes important. Data warehouses are becoming more self-service, absolutely, are becoming more trusted data, but the question is, there's always an issue sometimes with real timeness of data, right? So data warehouses traditionally have been dealing with batch oriented data, right?

Is it a data which is like one second away from my transactional data? The answer is no, right? There's always some delay there, but now with data pipelinings emerging, it can narrow the gap down, right, because you move data quicker in a timely manner into these warehouses, so we are starting to see data warehouses supporting these so to call near real time analytics, which I think is very important, especially with customers trying to embark that. So yeah, self-service, real time is starting to be addressed, and the end to end becomes critical with data warehouses as well.

But then there's also other means like data management for analytics that customers are doing, which can actually also help a lot as well, right? So yeah, yeah.

Okay, this one's a tricky question.

What your recommendations to ensure data security and protection and how do you test if it's effective? So how do you test security being effective?

Yeah, I mean, it is. What do you think, Mark? You probably know this better than me. So why don't you take this personally?

Well, my view is you should try and remove all the data before it, you know. That's about the gateway, That's my view, because I think it's very hard once it gets into a lake or a warehouse to really be secure, to test the security. And I haven't seen any good solutions to do it, to be honest, but I wondered if you knew of any.

No, I mean, you know, you catch the sensitive data beforehand, absolutely.

When it makes its way into the warehouses, it becomes so difficult to really start to look at it. No, there is people do a lot of checking of data, and, you know, one way of controlling that is sometimes the warehouse may have ten thousand tables, right, and you don't expose those ten thousand tables to all the people, guess, right? You only have fifty tables that have been authorized to view and use, number one. Hence, even though you have ten thousand tables coming in, that means nothing to anybody else because it's controlled, It's only fifty tables that matter, so when data comes in and it has nothing to do with those fifty tables, it goes into the other non fifty tables, you know, you're still protected, and now the important thing is that you want to control what data is going to be accessed by what business users and what have you, right, so that's one big thing.

Also, you have to actually have, what do you call, making sure your data is encrypted, data is masked, you gotta have transformation, you have got auditing, masking, vulnerability assessment, you got all those things at the platform as well. So you have to enable a lot of the security frameworks to enable these things. People forget about data warehouses, they all focus on transnational systems, right, this is where most of the effort is spent is on transnational system. They forget about data warehouses is also having a lot of sensitive data.

In fact, this is a question we think we asked about three years ago, like, how many of these enterprises are really protecting the data warehouses? And I think the result was like only seventy percent of the people who are protecting data, thirty percent were not doing anything about it, right? And that is a scary thing, which is scary, and if the hacker breaks in and tries to get into the warehouses, well, you're out of luck, because data is now stolen and no one has a clue about this actually. So, you know, and plus unauthorized users may get access to the data as well, but I agree with your point, you know, you really have to, as you ingest the data, those layers of the pipelines should determine what data gets into the warehouse initially from there, And that's the reason why, you know, the lakes have not opened up, right?

Why is the lakes so closely held? It's because the lakes could have any type of data, no one actually has a clue of that data. And then once the data, you know, that's why fifteen, twenty people only need people access their lakes, and in fact, one of the banks was speaking to me about this topic, and they're saying, we do plan to open up the lake in the coming years to hundreds of people, but we've to be very mindful about compliance and security as a result. But, you know, I think it's a very important point that people are going to take some gradual security measures, and big data never had a very strong security, whether it comes from encryption.

The main thing was not understanding the data context. When you dump data in a lake, you don't have a clue of what data is sitting in the lake, because there's no reference points. Unlike relational databases, you already have ten thousand tables. I know this ten thousand table only, nothing else.

But in the lake, it's all files. You don't have a clue about what files contain what data. So it's funny. Yeah.

Yeah, I agree. Agree.

Yeah. And even, I mean, I guess, you know, like I showed, field level encryption is a good way of actually just hiding and getting rid of some of that data. So you still have the data if you really need it, but it is protected, which is good pullback solution.

Yeah, that's actually, that's a good point actually. Well, what are the security controls Xpente has got for customers, I guess, of yours? How do customers of yours protect data in Xpente?

Oh, they do all those things that we've talked about, you know, masking data, removing data, encrypting data. We use field level encryption.

We've been partnering with Amazon to provide that.

So AWS KMS is a really nice solution because it means the customer has full control over the encryption.

So they can rotate the key when they want to rotate the key. They can actually enable us to encrypt the data, but not decrypt the data. And it's all using AES two fifty six bit encryption, which is really strong, which is great.

So we like it because it gives them full, you know, gives the customer full control over it. Yeah. And customers, you know, they go, as they're ingesting their data, they run a set of transformations depending on what the data is to either mask it or anonymize it or whatever they need to do as it goes through the platform. But yeah, we've seen a lot of customers very mindful now of GDPR. It seems like there's so many new standards coming out that everyone has to be really careful with sensitive data.

What about the keys management? Can customers do keys management as well? I know Amazon allows customers to actually do a keys management as opposed to Amazon doing keys management, right?

To do what, sir? The keys management, sorry, the keys.

Oh, yeah, the customer does that. Yeah, so in Amazon KMS, the customer actually gets full control over the keys. Yeah. And it's managed by, the actual key management is all managed securely by Amazon.

Not good. Yeah, yeah. Which is good. Key management is always one of the hardest bits of encryption.

It is a challenge.

Cool. Great.

Okay. And another question I have for you is with low code solutions and data automation, you mentioned a lot of that. Will that change the role of the data engineer? Because, you know, we mentioned that data engineers have been doing this for forty years now and the tools have changed.

We never had data engineers back in forty years ago. Remember that, right?

We had MIS engineers, maybe.

But we only had DBAs, you know, and develop and programmers, not even developers, got programmers, I guess, right, back then, thirty years, twenty, forty years I mean, you've been in industry for long enough, and so have I.

But, yeah, so your question is about what's the future of data engineers? Mean, or is it more about where do they go with low code, no code?

Is that the question?

Yeah, yeah, exactly. Like, I think the data engineer, as you mentioned, was really more like a developer. It's a programming role. You were writing in Python or Java, or probably Java in those days, or even something else.

And now with low code and no code, some of the data pipelining is actually being done by, you know, other people in the business, maybe business analysts or data scientists and things like that. And the pipelines are being done by more people in the company. And if you do transformations in your warehouse, you know, the sequel is being written by more people in the company. So I just wondered how that data engineering role is shifting in this new world.

Yeah, yeah. You know, I mean, well, first of all, you know, engineers are doing more with less in the models that, you know, new use cases are emerging, right? I mean, you know, remember a decade, two decades ago, we were just doing what in IT? Building CRM and ERP systems only. Nothing else, nothing else was going on, right? Think about where we are today, right? Are going new use cases every other day, every week, every day sometimes, you know, mean, so there are different variations of this, right?

So there's a lot of things happening with this, and data engineers are keeping themselves busy with all these things that need to innovate, right? So it's a very, very different world today in IT organizations, where they need to actually be able to do more of these use cases, more of these business initiatives, which were different than a decade, two decades ago. That's one big thing, right? And so data engineers are doing a lot of more new development around this data pipe plannings for these newer use cases, as I mentioned, which is obviously requiring skills.

Now having said that, you're right, I mean, is obviously a trend towards business users, business analysts automating, simplifying that pipelines and for them, but they are not the guys who are actually building those comprehensive end to end customer three sixty, not today, maybe ten years from now, five years from now, we will probably get there, but you know, today, they're actually solving this line of businesses, you know, in fact, it is funny, one of my one of the companies I was speaking to recently, they had a these the departments running very large spreadsheets, they're running like fifty gigabyte spreadsheets.

And like, woah, like, are guys doing a spreadsheet so large, I guess, why don't you put this into like a data mart or something like that, right? So, so, so all these smaller cases are emerging, this is where businesses are spending more time in building those pipelines for these use cases for these departments and line of businesses in operation, and that's where we're seeing most of these efforts being spent. But you know, so that's helping that, that's one thing, But the engineers that we're seeing are doing a lot more than what they are doing over the last couple of years, just because the fact there's a lot of these newer generation of building of these frameworks and use cases which are emerging, which is critical, right, especially with edge computing, right, edge computing, like driverless cars, and edge computing on these different edges for different devices and what have you, which is changing the industries altogether.

Those also are requiring data movement, by the way, right? I mean, how do you move data move data movement from those travelers cars into your data center? Right? Think about that, right?

So so that also, so all these pipelines are being built a call across, whether it's from the edges to the cloud to multi cloud to hybrid cloud, to legacy environments, towards the new generation environments, that really changes everything, right? So so, so I would say data engineers have been keeping busy. And I think I think they want automation, because they can do more with less what it means that they can do all this newer generation of pipelines, and make them available rapidly as well. You know, data engineers are not just about data pipelining, but they also do data warehouse work, data lake, sometimes data lake building as well, data supporting of that.

They also do schema structure building, schema definitions as well. So they're doing a variety of work in New Orleans as well, right? And typically, what are background of data engineers, they are mostly DBAs actually, have, because databases are becoming very highly automated, so they are kind of changing some of these profiles or roles towards becoming a data engineer, which is a more broader perspective of data lake, data warehouses, data pipelining, data integration, data quality, master data management, providing that framework towards dealing with all of these integration points, which makes it a lot appealing for customers as well.

Customer, mean, for the business itself. So certainly, this is a very important direction. We see data engineers putting a lot of kind of work, and if you look at LinkedIn, you'll see a lot of new new job jobs around data engineering, lots lots of them. Why?

Because this is a gap. This is a gap because we don't really have that persona in organisations, we don't really have roles. We don't have expertise in this. And this is a big area.

And that's why a lot of people are hiring in retailers and financial services and healthcare, go to look at LinkedIn, and you find out there's tons of jobs in this area.

Okay, great. Well, you. Thanks. It's a great answer.

I think we've run out of time unfortunately, but thank you very much. So thank you everyone for attending today's webinar. I'd like to personally thank Noel for his insights on modern ETL. It's been great.

And let me restate the bottom line from my talk. ETL is still very relevant to clean process and secure your data. And ETLT gives you the best of both worlds. It gives you ELT speed and flexibility and protection of your sensitive data.

So that ends today's session on optimizing your data strategy with modern ETL. And please contact us if you'd like to learn more or if you have more questions. So thank you very much and thank you again, Noel.

Thanks.