Last night I saw a feature on the making of Avatar. What caught my attention about this was the way in which they had created a 'virtual camera' through which James Cameron could observe the live action of his actors morphed into their avatar bodies and displayed in the CGI generated world of Pandora - in REAL-TIME.
In my business world, which is about as far removed from Pandora as you could get, I build data warehouses and have been doing so for many a year. From time to time the concept of Real-Time Data Warehousing has arisen and largely been dismissed for the following technical reasons:
- Data Dependencies
- Slowly Changing Dimenions
- Complex Transformations
- Updating Aggregate/Summary Data and Cubes
All of the above are still valid, and there would still need to be a valid business reason to go to the extra cost and expense of building a real-time data warehouse as opposed to a cheaper batch one.
However, I'm sure the technical hurdles that the Avatar movie makers overcame must have been a whole lot longer than my sorry list. Maybe it's time for a rethink!
Showing posts with label BI. Show all posts
Showing posts with label BI. Show all posts
Sunday, February 7, 2010
Tuesday, February 2, 2010
ROLAP smolap
One of the first ROLAP tools that I came across was Oracle's Discoverer product. As one of Larry's consultants I lead a Data Warehouse Team that delivered our reports using it when it was a brand new product. So new in fact that the client didn't realise that the paint hadn't dried on it and it was actually pre production software. They assumed that Discoverer 3.0 had been preceeded by versions 1.0 and 2.0. and there's another story in there about trusting Oracle Sales and Marketing, but I digress.
Some 12 years later I came across Oracle Discoverer again. To my suprise very little appears to have changed. The EUL and full client looked almost identical. I'm sure that under the hood there have been some changes for intranet, pdf's and web delivery but I'm still a bit amazed about the lack of innovation in the ROLAP world.
Business Objects finally seem to be getting things together with BOXI R3 and I have to admit that I haven't seen Cognos's stuff for a while so for them I can't comment.
The only real innovation I've seen in the last few years was ProClarity, before they were swallowed up by Microsoft, but that's OLAP and not ROLAP offering.
A few years ago I started to believe that the Reporting tools were stagnating and that cubes - whether OLAP or ROLAP weren't the answer. I hoped that the move into RIA (Rich Internet Applications) and tools like Curl would fill that gap but as yet nothing seems to have developed there.
Maybe new platforms like the iPhone and more importantly the iPad will spur the sleeping Reporting giants into a new series of innovation. I hope so.
Some 12 years later I came across Oracle Discoverer again. To my suprise very little appears to have changed. The EUL and full client looked almost identical. I'm sure that under the hood there have been some changes for intranet, pdf's and web delivery but I'm still a bit amazed about the lack of innovation in the ROLAP world.
Business Objects finally seem to be getting things together with BOXI R3 and I have to admit that I haven't seen Cognos's stuff for a while so for them I can't comment.
The only real innovation I've seen in the last few years was ProClarity, before they were swallowed up by Microsoft, but that's OLAP and not ROLAP offering.
A few years ago I started to believe that the Reporting tools were stagnating and that cubes - whether OLAP or ROLAP weren't the answer. I hoped that the move into RIA (Rich Internet Applications) and tools like Curl would fill that gap but as yet nothing seems to have developed there.
Maybe new platforms like the iPhone and more importantly the iPad will spur the sleeping Reporting giants into a new series of innovation. I hope so.
Labels:
BI,
Business Objects,
Curl,
Data Warehousing,
iPad,
iPhone,
Microsoft,
OLAP,
Oracle,
ROLAP
My Online Doppelgänger
It seems that a few months after I started my IT Journeyman blog that I have an online doppelgänger. That's OK because I'm not the jealous type.
Having skimmed said blog the following post http://www.itjourneyman.com/2010/01/16/data-warehouse-2nd-time-is-a-charm caught my interest and it's essentially a rehash of a few white papers on "pitfalls/mistakes to avoid when building data warehouses". The long and the short of the post is that your first data warehouse will be a failure but don't worry because the second one will learn from those lessons and succeed.
I love to say that this was true but imho its just not that simple. In my travels I've worked on first stab data warehouses that have been blinding successes and also third tries that have had no more luck than their predecessors.
There are lots of elements that go into making a data warehouse project succeed or fail and often the initial expectation setting exercise is crucial. We have to be very careful in determining the criteria of what makes a data warehouse work and what doesn't.
It's a bit like marriage and divorce. Most people would assume that definition a fifty year marriage must have succeeded - but what if the husband and wife were at each others throats for the duration. Likewise divorce after 10 years is seen as failure but what if you've produced a couple of wonderful and well adjusted kids and went your own way amicably. Expectation is everything.
What I can say is that in my experience Data Warehouse projects are difficult and that's why I choose to work in that field and not implemeting somebody elses off the shelf package.
Data Warehouse Projects are voyages of discovery and it's what we learn along the way and not necesarily where we end up that's really important. The problem is that most organisations and most PM's just don't understand that yet.
Having skimmed said blog the following post http://www.itjourneyman.com/2010/01/16/data-warehouse-2nd-time-is-a-charm caught my interest and it's essentially a rehash of a few white papers on "pitfalls/mistakes to avoid when building data warehouses". The long and the short of the post is that your first data warehouse will be a failure but don't worry because the second one will learn from those lessons and succeed.
I love to say that this was true but imho its just not that simple. In my travels I've worked on first stab data warehouses that have been blinding successes and also third tries that have had no more luck than their predecessors.
There are lots of elements that go into making a data warehouse project succeed or fail and often the initial expectation setting exercise is crucial. We have to be very careful in determining the criteria of what makes a data warehouse work and what doesn't.
It's a bit like marriage and divorce. Most people would assume that definition a fifty year marriage must have succeeded - but what if the husband and wife were at each others throats for the duration. Likewise divorce after 10 years is seen as failure but what if you've produced a couple of wonderful and well adjusted kids and went your own way amicably. Expectation is everything.
What I can say is that in my experience Data Warehouse projects are difficult and that's why I choose to work in that field and not implemeting somebody elses off the shelf package.
Data Warehouse Projects are voyages of discovery and it's what we learn along the way and not necesarily where we end up that's really important. The problem is that most organisations and most PM's just don't understand that yet.
Labels:
BI,
Data Warehousing,
IT,
Project Manager,
Software Development
Monday, February 1, 2010
Assisting the Police with their inquiries
Back in 1998 I was doing some Pre-Sales Consulting for an Account Manager trying to sell a Data Warehouse solution to a local state police force. I badgered the salesman to let me use the above title as a tagline on the demo but unsurprisingly he didn't see the funny side.
During the demo the thorny question of Metadata came up. More precisely - Consolidated Metedata. As I'd just come off a project where I'd defined the Metadata Architecture and Solution I was well qualified to answer the query.
At the time we had three sources of metadata for our solution. These were:
- The Database Data Dictionary
- The CASE/Data Modeling tool in use
- The ROLAP Semantic Layer
Note that this we didn't use an ETL product that would have been a fourth source of Metadata.
Now the interesting thing here is that all the software was written by the same company in the same software labs so one would hope that some level of shared metadata would be possible. Alas no. Not only did the metadata in each repository overlap but there was no easy way of combining it into a single source of consolidated metadata repository.
I answered the question honestly that nobody had a good story here, not us nor our competition. I think the client appreciated my honesty here. The account manager obviously not wanting to leave a bad impression did what all account managers are prone to do and started promising vaporware with some cock and bull story about the software labs in California working on that problem.
The interesting thing is that here we are over a decade later and I've still to see a good answer to this problem.
During the demo the thorny question of Metadata came up. More precisely - Consolidated Metedata. As I'd just come off a project where I'd defined the Metadata Architecture and Solution I was well qualified to answer the query.
At the time we had three sources of metadata for our solution. These were:
- The Database Data Dictionary
- The CASE/Data Modeling tool in use
- The ROLAP Semantic Layer
Note that this we didn't use an ETL product that would have been a fourth source of Metadata.
Now the interesting thing here is that all the software was written by the same company in the same software labs so one would hope that some level of shared metadata would be possible. Alas no. Not only did the metadata in each repository overlap but there was no easy way of combining it into a single source of consolidated metadata repository.
I answered the question honestly that nobody had a good story here, not us nor our competition. I think the client appreciated my honesty here. The account manager obviously not wanting to leave a bad impression did what all account managers are prone to do and started promising vaporware with some cock and bull story about the software labs in California working on that problem.
The interesting thing is that here we are over a decade later and I've still to see a good answer to this problem.
Taking the Mountain to Mohammed
I've been working in the field of Data Warehousing for some 13 years now. Actually my first every data warehouse was a Reporting System I did back in 1992 long before I'd ever heard the terms DW & BI but that's another story.
The interesting thing that, so far, has been a constant in all that time, no matter what style of Data Warehouse (from full blown Inmon Corporate Information Factory to Kimball Federated Data Marts), is that we extract data from source systems and move it and load it into a data warehouse (be it an EDW, Data Mart, ODS, RDS, whatever). We'll use terminology like ETL, OLAP, ROLAP, Cubes, Star Schemas, Metadata, Slowly Changing Dimensions, etc. along the way to baffle the business and make ourselves seem clever but fundamentally any data warehouse or data mart involves moving data from a source system into target reporting system.
Back in the 90's this made perfect sense because it was inconceivable that we could slap resource consuming queries on reports against the mission critical core business systems.
Nowadays that just not the case. There are many technical solutions out there that could enable us to place a large and significant batch query and reporting load against our production data that would have zero impact on the core business systems. Technologies that spring to mind include Server Virtualisation, Disk Replication and Mirroring, O/S and Database Parallel Server technologies, etc.
The question is why don't we employ these technologies? I suspect that in the field of DW & BI we're in a stuck in a Kimball or Inmon rut and that for the time being we will continue to Take the Mountain to Mohammed.
Ah, but what about history I hear you ask? Well yes it's true that we often capture history in the data warehouse that we cannot keep in our online systems but often the need and justification for history is overstated. Besides another way in which we could keep all the history we'd ever need (and we probably already do this to some degree anyway) is to ensure that all PDF reports that are produced are kept online in some fashion. There are alternatives if we are creative.
Maybe within the decade well see a shift away from this and let Mohammed walk to the mountain for a change.
The interesting thing that, so far, has been a constant in all that time, no matter what style of Data Warehouse (from full blown Inmon Corporate Information Factory to Kimball Federated Data Marts), is that we extract data from source systems and move it and load it into a data warehouse (be it an EDW, Data Mart, ODS, RDS, whatever). We'll use terminology like ETL, OLAP, ROLAP, Cubes, Star Schemas, Metadata, Slowly Changing Dimensions, etc. along the way to baffle the business and make ourselves seem clever but fundamentally any data warehouse or data mart involves moving data from a source system into target reporting system.
Back in the 90's this made perfect sense because it was inconceivable that we could slap resource consuming queries on reports against the mission critical core business systems.
Nowadays that just not the case. There are many technical solutions out there that could enable us to place a large and significant batch query and reporting load against our production data that would have zero impact on the core business systems. Technologies that spring to mind include Server Virtualisation, Disk Replication and Mirroring, O/S and Database Parallel Server technologies, etc.
The question is why don't we employ these technologies? I suspect that in the field of DW & BI we're in a stuck in a Kimball or Inmon rut and that for the time being we will continue to Take the Mountain to Mohammed.
Ah, but what about history I hear you ask? Well yes it's true that we often capture history in the data warehouse that we cannot keep in our online systems but often the need and justification for history is overstated. Besides another way in which we could keep all the history we'd ever need (and we probably already do this to some degree anyway) is to ensure that all PDF reports that are produced are kept online in some fashion. There are alternatives if we are creative.
Maybe within the decade well see a shift away from this and let Mohammed walk to the mountain for a change.
Subscribe to:
Posts (Atom)