Showing posts with label Big Data. Show all posts
Showing posts with label Big Data. Show all posts

Tuesday, February 14, 2012

Unstructured data is a myth

Couldn't resist that headline! But seriously, if you peel the proverbial onion enough, you will see that the lack of tools to discover / analyze the structure of that data is the truth behind the opaqueness that is implied by calling the data "unstructured".

The need to take a deeper look at this? See this graph:
A lot of data growth is happening around these so-called unstructured data types. Enterprises which manage to automate the collection, organization and analysis of these data types, will derive competitive advantage.

Every data element does mean something, though what it means may not always be relevant for you. Let me explain with common data sets which are currently labeled "unstructured".

  • Text: Lets start with the subsets in here. 
    • Machine generated data (sensors, etc) definitely can be deciphered once you get the meta data structures / templates that the machine uses to generate the data. Of course, some of the fields in the stream will need more advanced analysis/discovery capabilities to automate the analysis.
    • Interaction Data: This is the case for social media data where a lot of business value lies in the long open text fields where people express sentiment about other people and products. To automate the analysis of these, entity recognition and semantic analysis provide the ability to understand the data better. In other words, if you can represent the text data as a collection of entities, relationships between them and relationship attributes like sentiment, you are much closer to analyze the data than you might think!
  • Images: Image recognition algorithms have almost become mainstream (though not very well-received as seen in the reservations against Google and Facebook deploying these at scale). Again, these techniques yield entities though deriving relationships and sentiment are much more challenging.
  • Audio: Again a lot of research is yielding technology which can decipher the content of audio streams and even annotate the resultant content with mood of the speaker! You could then leverage the text analysis techniques to get closer to the analyzable data.
  • Video: Unarguably, this is the most challenging data type due to the sheer volume of data that needs to be handled. Image recognition techniques can be applied per frame or a series of frames to extract entities. Of course, deciphering the action (the video content) is further out in the future. Audio recognition can be applied to understand part of the "action" content.
Based on the above, some new data handling and analysis capabilities are required to extract more value out of these new data types.
  • Dynamic Meta data discovery: This is mainly for text data. This includes the ability to
    • Dynamically derive meta data out of sample result sets e.g. new REST end points
    • Maintain / Master metadata on an ongoing basis
    • At run time, choose the appropriate / best matching metadata set out of several possible options
  • Taxonomy Setup: You need to be able to capture / represent your business and its entities for other analysis layers to reference and annotate incoming data. As your business evolves, this taxonomy will get richer.
  • Entity Extraction and Semantic Analysis: This provides the ability to apply the taxonomy to any text data stream and derive entities and relationships expressed in that stream. This analysis can then be stored either in a relational database or as a graph.
  • Multimedia Recognition Techniques: As described earlier, various techniques for deciphering the content of images, audio and video are required to analyze these data types.
The layering is along the following lines:

A lot of action is still on the top layers but eventually it will encompass audio and video as well.

Do you still believe all of this data deserves the opaque sounding "unstructured" tag? Are you building the capabilities to put the structure back into this data?

_______________________________________________________________________________
Ram Subramanyam Gopalan - Product Management at Informatica
My LinkedIn profile | Follow me on Twitter
Views expressed here are personal and do not necessarily represent those of Informatica.
_______________________________________________________________________________

Monday, January 30, 2012

How would your enterprise's social graph look like?

Imagine having a Facebook-like rich social graph tailor-made for your company with just the right information about  relevant entities and their activities being captured and maintained. You can derive accurate insights about your customers; excel in marketing with highly segmented campaigns using personalized content; get better returns on marketing by focusing on influencers; convert more of the resulting leads with personalized offers and strong referrals; get the best return on support by focusing on the key (vociferous?) customer segments; improve key product dimensions by analyzing feedback; keep a tab on competitors in the context of your most profitable products; ... If you are getting ideas, read on!

Before you do anything else, convince yourself that you might be unknowingly taking undue risk by hinging your enterprise wagon entirely on data from third-party social networks - here are a few factors you might want to consider.

Assuming you are at least ready to weigh the benefits and costs of building your own enterprise social graph, here are some seeds for your thought process. 

What would an Enterprise Social Graph (ESG) look like?
The ESG is a highly interrelated graph that comprises 
  1. Entities like people (customers, employees), companies (the enterprise itself, competitors, partners, suppliers), products (those owned by the enterprise and its competitors)
  2. Defined Relationships among these entities
  3. Activities with one or more entities as actors and/or subjects - Documents can represent these activities
Though the ESG does not really have layers or strata, it would be useful to visualize it as a layered graph with connections running between nodes in the same layer as well as across layers.

Starting with the Enterprise Graph
If you represented the relationships between entities that can be gleaned from enterprise systems (not knowledge hiding in employees" minds!) as a graph, it might look something like this:



You might have noted the following:

  • Enterprise Entities:The first layer is illustrative of the kinds of entities that are typically seen within the enterprise. The taxonomy of the entities will be industry-specific and the canvas is as large as it can get, with technologies, partners/suppliers and competitors all coming into play.
  • Activities: The kinds of actvities that are captured correspond to one of those in CRM, HR, PLM and ERP systems. You can see that this already catches relationships like "Customer A owns Product P", "Customer B complained about product Q", "Employee E manages Employee F", "Employee G knows Influencer I", etc.
  • People: Several links between people are missing in this and that is deliberate. e.g. relationships between influencers and employees/competitors. You might argue that someone in the company "knows" about this but the counter to that is if you ran an algorithm to weigh the influence of that person, we would miss out on this fact! In general, as you move further away from your employees and customer entities, you have lesser information on people and activities available for analysis.
  • Mostly the relationships captured in this version of the graph (no social dimension to the entities) are less dynamic (e.g. something more defined like a Product-Technology hierarchy). Many enterprises do not have the capabilities to process data such as Survey free-form fields, chat transcripts, etc which are captured and stored in the enterprise systems.

Adding the Social to the Enterprise Graph
Lets add the social dimension which means we incorporate the following additional information that we can glean from social media monitoring:
  • People's likes and dislikes; skills, preferences
  • People's personal and professional connections
  • Social Media activities of people and companies around technologies and products of your interest
  • Competitive moves and directions
  • Motives, Intents and Drivers for people's buying behavior
  • and so on ...

Adding just a few illustrative examples from the above list, the ESG might now look like this:

As you can see, this graph holds information for use within the enterprise as well as the external world. You can start to see missing links in your corporate puzzle that might explain several business trends that you see in your traditional BI systems but can't find reasonable causes and remedies for!

How to build and use the ESG?
You can imagine this to grow immensely dense for large enterprises or even for small enterprises which want to include more information in this graph. Beyond a point, the scale of this graph will necessitate Big Data capable solutions. You can of course start small and build this graph over time but eventually you should plan for large scale graphs especially if you want to (and you should) do temporal analysis of trends.

Looking at the building blocks for a social data integration solution, the ones that will be key for building the ESG are the Text Analytics modules and the infrastructure components to reflect the "graph" nature of the information. Topic for a joint follow-up post from our dev team and me!

As a parting thought, there are several vendors with patents around this area, but not very many published success stories of enterprises building powerful social graphs. (If you know of any, please do leave a link or two in the comments or tweet them to @ramsgopa).

There are two ways to react to that fact - either wait for other enterprises to taste success (and then follow them) or be one of them. How will you react?


_______________________________________________________________________________
Ram Subramanyam Gopalan - Product Management at Informatica
My LinkedIn profile | Follow me on Twitter
Views expressed here are personal and do not necessarily represent those of Informatica.
_______________________________________________________________________________

Thursday, January 26, 2012

Do not depend on Facebook (alone) for customer insights

Diving deeper into one of the key areas where all enterprises can find value from integrating social media, customer profile enrichment with the social dimension can help several business functions.

Even though I have been thinking that companies need to maintain their own customer profiles, I got emboldened to take a stronger stance, by this white paper (Regn./PDF) from @gxsoftware on a related subject of avoiding over-dependence on external social networks like Facebook to gain deeper insights on their customers.

Though my headline has Facebook, this applies to all of external social media sources.

Factors to consider
Here are the factors you should consider in your customer profile building strategy:
  1. Your data needs may not be easily satisfied: Though the social networks do try and expose as much data as possible through APIs and Aggregators, you might not get the data in the most efficient manner for your specific needs. There are myriad restrictions on APIs (data volumes, number of calls at various levels like IP Address, App, API and User) which need to be negotiated through to get the data. Also, not every type of activity on the social network might be of interest to you.
  2. Privacy policies unclear and likely to toughen up: As Ray Wang lays out demands in his recent post, users are likely to press for better privacy controls on their favorite destinations like Facebook. In fact, there are startups offering fine grained user-driven privacy settings e.g. Diaspora. So you should realistically expect unpleasant surprises on non-availability of data that you assumed will be available forever!
  3. Data Ownership is also unclear: Even if you had the data available, the social media that hosts the data and/or the end-users may assert uncomfortable-sounding rights on the data and its permitted usage/storage etc. This may lead to expensive workarounds/solutions to adhere to these rules if you haven't taken care to design it upfront.
  4. Data access may become pricey: Finally, even as you work through these issues, you might end up staring at steep fees to access the data. After all, Facebook and the like will monetize the data in as many avenues as they can (advertising, data access, etc). 
The WP lists a few other factors like losing competitive advantage since everyone can get the same data access - these are valid factors too but they are almost a given in the social world.

What can you do?
As you might have realized, many of the factors are beyond the enterprise's control. So the enterprise should have a plan of action to mitigate the downsides. Some of the elements in that plan could be:
  1. Start Listening now: Even as the privacy and data ownership rules continue to toughen up (and they should), you should be collecting as much data as possible (while of course still adhering to the T&C's of the social media!) by engaging with customers and listening to their activity streams. Start your social media integration journey now.
  2. Look at a policy-driven enterprise social platform: As you start listening, remember the value of an enterprise-wide platform leveraged by multiple business functions for multiple use cases. One of the first issues you will have to solve is to gain connectivity to the social media (see factor#1). In a subsequent post, I will describe the challenges in this area and what to look for in a possible solution. The platform should ideally have policy-driven (and automated) models pertaining to data ownership, access rights and storage restrictions e.g. you should be able to dynamically model a source as having members-only access with data storage restrictions on certain fields in the incoming data. The platform should be able to enforce the data access and data archival based on these settings. At scale, poilcy-driven automation is key.
  3. Create your own Enterprise Social Graph: We shouldn't expect others to do the heavy lifting for our business needs. You should evaluate how you can mashup enterprise and social dimensions for your key entities like Customers, Products and Suppliers. If you already have a MDM-based Customer/Product Hub, it would provide an ideal platform for this. What you might end up with is a logical graph which captures the dense relationships between and among these entities. Ideally, all your business cases should be able to query this graph to get the required information. This is truly Big Data territory - a salivating possibility for another post!
  4. Invest in community building: Now that your customers see enough value in engaging with you across multiple channels (mainly social media since it is ubiquitous), you should be able to invite them to participate in your own communities or at least on forums where you have access and maybe even ownership of all the data. The data access terms are now more favorable to you.
  5. Leverage this graph in CEM and PLM: The Enterprise Social Graph can now be even more relevant and enriched with very detailed customer profile information. The rich profiles that this graph provides for your customers and products is ripe for leveraging in improving the customer experience and your products.
What is your take on how the data access situation will play out in the medium-term? Where are you storing / planning to store the social dimension of your customers and products?


_______________________________________________________________________________
Ram Subramanyam Gopalan - Product Management at Informatica
My LinkedIn profile | Follow me on Twitter
Views expressed here are personal and do not necessarily represent those of Informatica.
_______________________________________________________________________________

Wednesday, January 11, 2012

Social Data Integration journey - Where are you on it?

There are several "Stages of Evolution" thoughts around how enterprises can go about ingesting social data into their business DNA. Adding to these, from my experience creating CRM and data integration products, here is an outside-in approach for enterprises towards social media.

Companies must first map the customer journey through various social media as they interact with the company. The following depicts such a typical journey (which, in most part, applies to both B2B and B2C businesses):

Note that some of these "Events / Triggers" are actually a cumulative experience for your customers.

Once this map is created, companies can decide on how to start their own journey. A phased approach is prudent especially in the social media arena where missteps can get amplified quickly. The following illustrates this:


Importantly, this is an additive approach in that companies will have to continue to listen and monitor on a broad set of social media while participating and innovating on the most effective among these social media streams.

How are you approaching this massive opportunity? Where is your company on this journey?

In subsequent posts, I'll share my thoughts on how enterprises can derive the most of their investments in social media.


_______________________________________________________________________________
Ram Subramanyam Gopalan - Product Management at Informatica
My LinkedIn profile | Follow me on Twitter
Views expressed here are personal and do not necessarily represent those of Informatica.
_______________________________________________________________________________

Monday, November 14, 2011

Big Data + Big Money = Big Change

The dust has barely settled on last week’s Hadoop World and the blogosphere is still abuzz with the implications of various announcements, analyst reports and miscellaneous provocations that are inevitable in events like this.   No better time than this blogger to jump into the fray! 

My takeaway is the evidence has never been stronger that Big Data is real, and several related announcements and data points bear this out.

Monday, November 7, 2011

Think End-to-End When Parsing and Visualizing Big Data in Hadoop

It has been a big, busy week at Informatica.  We introduced HParser, the industry's data parser for Hadoop. As part of the launch activities, we recorded a video discussion on parse and visualize Big Data, featuring Brett Sheppard from Zettaforce, Ronen Schwartz, Vice President of B2B Products at Informatica, and Karl Van den Bergh, Vice President, Product and Alliances at Jaspersoft. The discussion helps you understand the approaches to extend Jaspersoft business intelligence (BI) and Informatica data integration investments to leverage data stored in Hadoop. HParser is designed to reduce the need for manually written Map/Reduce scripts.  You can also check out the chalktalk and product demonstration of HParser .

Pace of innovation for Hadoop has been rapid.  One of the reasons that we are seeing this accelerated rate of growth in Hadoop is that the Hadoop community and vendors across the relevant data computing infrastructures seem very engaged. 

Friday, October 28, 2011

Free Your Data!


Much has been written in the blogosphere recently of Oracle's Public Cloud announcement at the recent Oracle OpenWorld, and what (according to Larry) it means for data portability and the integration industry in general.

Quoting Larry from his keynote, comparing with Salesforce.com:  "Our cloud's a little bit different ... our cloud is based on industry standards and supported full interoperability with other clouds. ... you can take any existing Oracle database you have and move it to our cloud.  Just move it across and it runs unchanged.  Oh by the way, you can move it back if you want to."  He went on to say that Force.com is the "roach model of cloud services" due to its use of custom/proprietary programming languages like APEX.  "You can check in, but you can't check out."




 This makes for entertaining press, but the subtext here is that all one needs for data integration is some open standard APIs and the ability to run an Oracle database instance anywhere.  In other words, it's all about portability.

Perhaps.  If you have chosen to delegate all of your data to an Oracle-only software stack, then knowing you are "free" to run that database pretty much anywhere you want may feel liberating, freeing you from the dark forces of vendor lockin.  In other words, avoid data lockin by ... wait for it ... consenting to lock in your data to one vendor's database?   Really? And this coming from one of the great industry consolidators of  the last decade, a strategy whose success depends on the maintenance revenue of customers keeping their data right where it is.

Monday, October 24, 2011

2012 Predictions Season already starting - Big Data front and center


2012 predictions season seems to be starting early this year, and not surprisingly Big Data is showing up on everybody's lists. The top 10 list from Nucleus Research caught my eye, not because of the obligatory Big Data listing (it showed up as #5), but because of how relevant Big Data is to several other trends that were listed.

As frequent readers of this blog will soon recognize as a common theme here - the real value in Big Data is in the opportunities it creates, far beyond solving its "big mess" management problems. In other words, maximizing the return on data.

Without going into the details of Nucleus' report (free download can be found here) specific areas that seem particularly ripe for Big Data related opportunity include:

Friday, October 21, 2011

The Real Value of Big Data

So you're an IT manager (or an architect, or a business analyst, ...), and a vendor tells you that the amount of data in your org is exploding, and unless you buy from them RIGHT NOW, then you'll be left with an unmanageable Big Mess. By now you have probably heard it often enough that you're not sure whether to believe it. You've gotten this far without the latest newfangled tool from company X, you're not yet up to your eyeballs in messes, so why the urgency? Why not just keep doing what you're doing?

Good question. Actually, any vendor that tries to sell this way is missing the point. I've heard the same pitches - enough times to wonder if they are all just copying the same script. "You're drowning in data, blah blah." Where's the imagination? Instead the point should be about the opportunities to be seized, not the problems to solve. It should be about the business value of that data - and helping you understand how you can benefit from it. For example:

Thursday, September 29, 2011

What is so Big about Big Data?

Welcome to the Big Data Integration Blogspot. What is Big Data Integration? Well, Big Data is the confluence of three technology trends:
  • Big Transaction Data: Massive growth of transaction data volumes
  • Big Interaction Data: Explosion of interaction data such as social media, sensor technologies, call detail records, and other sources
  • Big Data Processing: New very large scale processing with Hadoop and other alternatives
Big Data integration helps you fulfil the business potential of Big Data.  Well of course, you like to know what this means in the real world right?  Furthermore, Big Data can be Big in terms of business results but also in terms of risk. So in this blog, we are inviting technical professionals to share their ideas and approaches to Big Data.  We look forward to hearing your thoughts.