Saturday, 21 January 2023
Four starting points to transform your organization into a data-driven enterprise
Tuesday, 1 November 2022
Trustworthy AI helps provide equitable preventative care for diabetics
Healthcare organization uses technology to identify members for proactive care
What healthcare is doing with data fabric and AI to mitigate risks
How trustworthy AI provides better help
Wednesday, 14 September 2022
How to stay ahead of ever-evolving data privacy regulations
Adopting a privacy-centric approach built around a data fabric
Build a foundation using a common catalog and metadata
Operationalize data privacy through automation
Govern data and allow self-service consumption
Tuesday, 26 July 2022
Data fabric marketplace: The heart of data economy
In older civilizations, where transportation and communication were primitive, the marketplace was where people came to buy and sell products. This was the only way to know what was on offer and who needed it. Modern-day enterprises face a similar situation regarding data assets. On one side there is a need for data. Businesses ask: “Do we have this kind of data in the enterprise?” “How do we get that data?” “Can I trust that data?” On other side, enterprises and organizations sit on piles of data, and they have no clue that others need it and are ready to pay. Like a medieval marketplace, a data marketplace can bring these two sides together to trade.
This discussion is more relevant with the advent of data fabric. Data fabric is a distributed heterogeneous architecture that makes data available in the right shape, at the right time and place. A data marketplace is often the first step toward the data fabric vision of an enterprise. A data marketplace tops our major clients’ wish lists, a trend also observed by industry analysts. For example, Deloitte identified data sharing made easy as one of the top seven technology trends. Gartner predicts that by 2023, organizations that promote data sharing will outperform their peers in most business metrics.
Why is marketplace the centerpiece of data fabric?
The main purpose of data fabric is to make data sharing easier. Today, when data sharing roughly equates to data copying, enterprises spend a lot to move data from one place to another and to curate the data to make it fit for purpose. This long journey of data discovery and processing can lengthen the application development lifecycle or delay insight delivery. Enterprises must address the inefficiencies to remain competitive. Business users and decision makers should be able to discover and explore the data by themselves to perform their jobs through self-service capabilities. Enterprises want a platform where data providers and consumers can exchange data as a commodity using a common and consistent set of metadata. Doesn’t that sound very similar to the marketplace model?
How does a marketplace make it happen?
To make data sharing an integral part of the culture, the data governance practice of an organization must associate certain measurable KPIs against it. Those KPIs can be met through incentivization schemes. So, the marketplace must have some monetization policy defined for the data, even for internal sharing. (The currency may not always be money. Reward points can also serve the purpose.)
From the technical perspective, a marketplace depends on two capabilities: a strong foundation of metadata and the capability to virtualize or materialize data. The metadata creates a data catalogue similar to the product catalogue in any typical e-commerce platform. This allows data providers to publish their data products to the platform with appropriate levels of detail (including functional and non-functional SLAs), where data consumers can discover them easily. Data virtualization or materialization capabilities also help to reduce the cost of data movement.
Data marketplace vs. e-commerce platform
A data marketplace has a few differences from an e-commerce platform. Most e-commerce platforms are either marketplace-based or inventory-based. But a hybrid approach is essential for a data marketplace. Individual organization units of the enterprise will offer their respective data products. But some data products should be owned and offered at the enterprise level. This is because some of the datasets (e.g., master or reference data) may need to be cleansed, standardized and de-duplicated from multiple sources to offer a single view of truth across the enterprise.
For example, a financial institution may have several lines of business (LoB) such as banking, wealth management, loans and deposit. A single customer may have presence in all four LoBs, causing separate footprints. When the analytics department wants to get a 360-degree view of the customer to run an integrated campaign, customer data is integrated into one place to generate a single value of truth. The marketplace can be this place of consolidation. In these cases, the marketplace may have to maintain its own inventory of data — thereby adopting a hybrid approach.
The second difference between a data marketplace and a typical e-commerce platform is the nature of the product. Unlike any typical product of e-commerce, a data product is non-rival in nature, meaning the same product can be provisioned for multiple consumers. The provisioning of data follows certain data rules as defined in the policy of the concerned dataset. So, data as a service would involve the hidden complexities of creating dynamic subsets (on-the-fly or cached) that are transparent to the consumer.
The third difference is the desired marketplace experience. The consumers of a data marketplace would like to explore the available datasets before procurement. This exploration is much deeper than a “preview” of the product that is typically available on e-commerce platforms. This means the marketplace should integrate with some development environment (such as Jupyter notebook) for better data exploration.
The fourth and final difference, which might be available only in a matured data fabric, is the capability to aggregate the data. Data marketplace consumers should be able to make intuitive queries that can be resolved through a synthesis of multiple data sources. This requires a highly illustrated business and technical metadata which form a knowledge graph to resolve such intuitive or semantic queries.
At present there is no single product in the market that provides all the typical e-commerce platform features and fulfills all of these requirements. But there are players who provide subsets of the features. For example, most market players have improved their capabilities in data cataloging, and there is increased interest on the client side to properly define their enterprise data sets to ease classification, discovery, collaboration, quality management and more. We have seen data cataloging interest grow from 53% to 66% within a year.
The Watson Knowledge Catalog, available within the Cloud Pak for Data suite, is one of the most powerful products in the cataloging space. On the other hand, Snowflake’s Data Marketplace and Exchange and Google’s Dataplex are ahead of the curve in providing access to external data in a pure marketplace model. The data marketplace of today would likely be a combination of many products.
Where can you start?
Data marketplace is likely to go through various maturity cycles within each organization. It can begin as a catalog of data products available from multiple data sources. Then the marketplace owner can create a few foundational capabilities. For example, clients would need business and technical metadata to define, describe, classify and categorize the data products. The users of the catalog would also be able to associate governance policies and rules to control access to the data for the intended recipients, which could be reused when appropriate data provisioning workflow is in place. At a later stage, marketplace features can be added to the catalog to publish internal and external data for the consumer to provision through a self-service channel. Once that is done, the marketplace can further mature to become a full-scale platform that facilitates data exploration, contract negotiation, governance and monitoring.
Source: ibm.com
Saturday, 16 July 2022
Do you know your data’s complete story?
Data is everywhere in a hybrid and multi-cloud world. Enterprises now have more data, more data tools, and more people involved in data consumption. This data proliferation has made it harder than ever before to trust your data: knowing where it came from, how it has changed, and who is using it. Data provenance is a complexity facing many clients engaged in data governance-related use cases. To help our clients overcome these challenges, we are pleased to announce our collaboration with MANTA to bring MANTA Automated Data Lineage for IBM Cloud Pak for Data to market.
MANTA Automated Data Lineage for IBM Cloud Pak for Data is a deep integration between MANTA’s end-to-end Data Lineage platform and Watson Knowledge Catalog on IBM Cloud Pak for Data. Data Lineage is an essential capability for modern data management and is a required aspect of regulatory compliance for many industries. Together with Watson Knowledge Catalog’s business friendly native data lineage, MANTA provides the most complete picture of technical, historical, and indirect data lineage.
How MANTA makes a difference
MANTA helps ease the amount of manual effort necessary for robust data lineage by providing scanners for the automated discovery of data flows in 3rd party tools such as Power BI, Tableau, and Snowflake. This information is then automatically scanned into Watson Knowledge Catalog’s Data Lineage UI and becomes available to view alongside the data quality, business terms, and other metadata previously available to Watson Knowledge Catalog users.
In addition to supporting the high-level summary view appropriate for many business users, clients can also dig deeper to see additional technical, historical, and indirect data lineage within MANTA’s Lineage Flow UI. Collectively this means that the addition of MANTA will provide quicker time to value not only through the automation of previously manual processes, but also through the ability to more rapidly answer questions about whether certain data is trustworthy.
A boost to your data fabric architecture
MANTA Automated Data Lineage for IBM Cloud Pak for Data will be available as an add-on to Watson Knowledge Catalog, further improving the ability of IBM’s data fabric solution to satisfy governance and privacy use cases. Surrounded by existing capabilities like consistent cataloging, automated metadata generation, automated governance, reporting and auditing assistance, MANTA Automated Data Lineage for IBM Cloud Pak for Data will help to bolster the data governance capability of IBM’s data fabric solution.
Of course, trust in data is important across every data fabric use case whether it happens to be building 360-degree views of customers or enabling trustworthy AI use cases. The multi-cloud data integration within the data fabric also helps connect the various data sources MANTA will be scanning. MANTA will simultaneously benefit and be benefited by the multiple data fabric entry points, that help customers on their data management strategy
What’s next?
The partnership with MANTA is just the beginning; we will continue to work closely to add more capabilities to MANTA Automated Data Lineage for IBM Cloud Pak for Data.
Source: ibm.com
Tuesday, 5 July 2022
5 recommendations to get your data strategy right
The rise of data strategy
There’s a renewed interest in reflecting on what can and should be done with data, how to accomplish those goals and how to check for data strategy alignment with business objectives. The amazing evolution of technology, cloud and analytics—and what it means for data use — changes quickly, which makes it easy to fall behind if your data strategy and related processes aren’t frequently revisited.
Read More: C2090-623: IBM Cognos Analytics Administrator V11
From multicloud and multidata to multiprocess and multitechnology, we live in a multi-everything landscape. Luckily, today’s data management approaches aren’t limited by traditional constraints like location or data patterns. The right data strategy and architecture allows users to access different types of data in different places — on-premises, on any public cloud or at the edge — in a self-service manner. With technologies like machine learning, artificial intelligence or IoT, the resulting insights are more sophisticated and valuable, especially when woven into your organization’s processes and workflows.
The evolution of a multi-everything landscape, and what that means for data strategy
As ecosystems transformed over the last few years and simultaneously increased the opportunities to improve results driven by data, a few main contributing factors drove major change in how you should think about your data strategy:
◉ The reality of hybrid multicloud and its accelerated adoption has created new possibilities and challenges. According to a recent Institute for Business Value (IBV) study, 97% of enterprises have either piloted, implemented or integrated cloud into their operations. But not all data is best suited for the cloud. While the share of IT spend dedicated to public cloud is expected to decline by 4% between 2020 and 2023, hybrid and multicloud spend is expected to increase up to 17%. Moving, managing and integrating data in a hybrid multicloud ecosystem requires the right data strategy, design and governance to eliminate silos and streamline data access.
◉ The diversity of data types, data processing, integration and consumption patterns used by organizations has grown exponentially. These data types require open platforms and flexible data architectures to ensure consistency with an appropriate orchestration across environments and strong re-approach of traditional capabilities and skillsets.
◉ The business areas need more value, faster — it’s a fact that the multi-everything landscape has triggered a more demanding world. Competition plays harder, and every day, new business models and alternatives driven by data and digitalization surface in almost every industry. Lines of business have increased pressure to speed go-to-market of innovation through new data-driven solutions, products or businesses. IT works to manage the underlying risk, security and performance through governance, without limiting flexibility. This balance between innovation and governance leads to new ways of working, like how the portability of data and analytics solutions has become a way to anticipate and adapt to change by enabling high flexibility to run in different environments and avoid vendor lock-in.
5 recommendations for a data strategy in the new multi-everything landscape
When it comes to getting a data strategy right, I like to apply some of the basic principles of a successful business model — scalability, cost-effectiveness and flexibility for change — and extend these concepts to technology, processes and organization. Organizations with data strategies that lack these factors often capture only a small percentage of the potential value of their data and can even increase costs without significant benefits.
In addition to the traditional data strategy considerations, such as recognizing data as a corporate asset or shifting to a data-driven culture with multi-functional teams, here are five recommendations for a data strategy that takes advantage of the multi-everything landscape:
1. Give data assets and accelerators top priority: Develop a process and culture around data that enables true standardization, re-use, portability, speed to action and risk reduction across the end-to-end data lifecycle. From the inception of use cases through the development, deployment, operation and scale of your assets, your data strategy should be supported by the right technologies and platform to enable fully operational and scalable solutions.
2. Establish a true enterprise-centric operational model: It’s critical to have the right operational model that’s fully aligned with the organization’s business objectives and its partnership ecosystem. This requires a deep understanding of the organization’s strengths and weaknesses. Embrace best practices but run away from pure academic approaches. Think big, but prioritize and articulate realistic and actionable plans, establishing the right partnership models along the way. That said, adopt and extend agile techniques as soon as you can.
3. Revisit the extent and approach for data governance: In this multi-everything landscape, data governance functions, processes and technologies should be constantly revisited to manage data quality, metadata, data cataloging, self-service data access, security and compliance across your enterprise-wide data and analytics lifecycle. Extend data governance to foster trust in your data by creating transparency, eliminating bias and ensuring explainability for data and insights fueled by machine learning and AI.
4. Don’t lose the basics: To improve business results, leverage data in a sustainable way and prioritize projects that are scalable, cost-effective, adaptable and repeatable to deliver both near- and long-term results. In all cases, the data strategy should be tightly aligned with your business objectives and strategy and built upon a solid and governed data architecture. It may be tempting to jump quickly into advanced analytics and AI use cases with the promise of astounding results without having considered every implication in the equation, but remember there is no AI without IA.
5. “Show and tell”: Take advantage of proven experiences, new technologies and existing assets as much as possible, and don’t forget to show results quickly. With the capabilities offered by hybrid multicloud environments and innovative co-creation and acceleration methods like the IBM Garage, you have the tools to design, implement and evolve your data strategy to continuously deliver on business outcomes. By showing tangible outcomes, fostering adoption and operationalizing at scale, you can reduce risk and accelerate the journey to a long-lasting, data-driven culture.
While the core principles of a data strategy remain the same, the ‘how’ has dramatically changed in the new data and analytics landscape, and the most successful organizations are the most adaptable to change when revisiting the data strategies. Today approaches and architectural patterns like data fabric and data mesh play an increasingly relevant role through enabling technologies and platforms like hybrid multicloud. As you look ahead, review your data strategy based on the opportunities presented in the new multi-everything landscape, and get ready for change.
Source: ibm.com
Thursday, 23 June 2022
Deliver on your data strategy with a data fabric
Today, organizations are experiencing relentless data growth spurred by the digital acceleration of the past two years. While this period presents a great opportunity for data management, it has also created phenomenal complexity as businesses take on hybrid and multicloud environments.
When it comes to selecting an architecture that complements and enhances your data strategy, a data fabric has become an increasingly hot topic among data leaders. This architectural approach unlocks business value by simplifying data access and facilitating self-service data consumption at scale.
At IBM’s recent Chief Data and Technology Officer Summit on data fabric and data strategy, I had an exciting conversation with data leaders from some of the world’s most prestigious organizations about how a data fabric architecture can help get the right data to the right people at the right time, driving better decision-making across the organization.
A data fabric orchestrates various data sources across a hybrid and multicloud landscape to provide business-ready data in support of analytics, AI and other applications. This powerful data management concept breaks down data silos, allowing for new opportunities to shape data governance and privacy, multicloud data integration, holistic 360-degree customer views and trustworthy AI, among other common industry use cases.
How IBM built its own data fabric
When I rejoined IBM in 2016, enterprise-level data and its use was having a pivotal moment. Because of advances in cloud computing and AI, it was clear that data could play a much bigger role beyond being a necessary output. Data was emerging as an asset that could benefit all aspects of an organization.
As the newly appointed Chief Data Officer (CDO), I was charged with creating a business data strategy built around making IBM a hybrid cloud and AI-driven enterprise. My goal was to implement a data strategy and architecture that gave appropriate users access to data. The key aspects were that the data had to be trusted and secure, and it had to deliver insights that drive business value through analytics without sacrificing privacy.
An important part of our evolution also included preparing for the EU’s General Data Protection Regulation (GDPR), which went into effect in May 2018. Our journey to build security, governance and compliance into our data strategy still serves as a digital solution for our clients and customers to achieve GDPR readiness.
Amid the ever-evolving complexity of our environment, we recognized a need for a data fabric architecture to deliver on our data strategy. The critical benefit of a data fabric is that it provides an augmented knowledge graph detailing where the data is, where it lies, what it’s about and who has access to it. Once we established an augmented knowledge graph, which is the main component of a data fabric, we were in a strong position to intelligently automate across the enterprise, infusing AI into all our major processes, from supply chain to procurement to quote-to-cash. This eventually delivered major reductions in cycle time. Plus, we soon realized that our own business transformation doubled as a blueprint for our clients and customers. But what else can a data fabric do?
How data fabric lays the foundation for data mesh
A data fabric not only acts as a central pane of glass that creates visibility; it also provides a flexible foundation for a component such as a data mesh. A data mesh breaks large enterprise data architectures into subsystems that can be managed by various teams.
By laying a data fabric foundation, organizations no longer have to move all their data to a single location or data store, nor do they have to take a completely decentralized approach. Instead, a data fabric architecture allows for a balance between what needs to be logically or physically decentralized and what needs to be centralized.
A data fabric sets the stage for data mesh in several ways, such as providing data owners with self-service and creation capabilities, including cataloging data assets, transforming assets into products and following federated governance policies.
The future of data leadership
Today’s data leaders are primarily focused on one of three strategic drivers: mitigating risk, growing the top line and enhancing the bottom line. What I find exciting about building a strategy around a data fabric architecture is that it allows data leaders to act as change agents by addressing these business needs all at the same time.
At the IBM summit, data leaders all concurred that we’re entering a new phase, where there will be much more decentralization in the world of data management. We expect to see concepts such as data fabric and data mesh play a critical role in strategy, as they can empower teams to access resources and tools they need on-demand to support them throughout the data product lifecycle.
But there’s more discussion to be had. In the newly-released guide for data leaders, The Data Differentiator, you’ll find our six-step approach for designing and implementing a data strategy. This information is continually tested and optimized by IBM experts during client engagements, and we wanted to share it with the community to help facilitate conversation about what it takes to succeed with data. You’ll also find a discussion of the role your data management architecture plays.
Source: ibm.com






