Wednesday, July 19, 2006

SaaS...

SaaS (Software as a Service). Now that’s the definitive word. I came across this story on SaaS today while going through the feeds. Over the period I have been hearing loads of terms: ASP, Software OnDemand and what not. And quite often have indulged into the lunch-time discussions around these topics. Oracle, SAP and other companies have OnDemand model for some time now. This big software companies helps enterprises reduce TCO and there IT cost in maintaining there systems by hosting them in their own server farms and managing, upgrading and patching them.

The next level for this evolving concept of Application Service Provider, Software OnDemand probably is naturally SaaS. During last couple of years I have heard about loats of software being delivered as service. Specially those related to content management and collaboration. Writely, Google Sheets, Salesforce.com, Basecamp, Sprouit.com and list goes on and on.

The article talks about nuances of SaaS and other related concepts like ASP and OnDemand. The marathon story covers what the big players like SAP, Oracle and Microsoft are doing to address SaaS in there next generations product line. The article covers how Salesforece.com fits in SaaS model while those offered by Oracle OnDemand and other vendors are far off from SaaS. The key thing which this article brings on table is the architectural principle of multitenancy, which means a single instance of the software runs on the provider’s servers, and all users log onto that same instance. The article goes on to talk about how SaaS cultivates a Web 2.0-like community. All and all good read and gives loads of insight on the SaaS landscape.

Tuesday, July 18, 2006

Ruby and RoR

Web application development has always fascinated me. The first nickel I earned after joining Information Technology Engineering was by developing a web application for my friend's uncle. In the initial days of my IT career I worked with startup Rightway Solution designing and developing small and medium sized web application. And man, those were exciting days of my career. That was the time of PHP. And ofcz PHP is a cool technology even now.

So developing Web based applications has excited me always. Few days back I was following some of the articles on trends of Web Development in 2005 and how the Web Development landscape's gonna look like in year 2006. And one of this article mentioned about R0R. Googling on RoR took me to this article from Curt Hibbs. The article gives you a head start on getting along with RoR. The RoR is a cool framework build on top of Ruby. And Ruby itself is a cool language. Damn powerful and as its mentioned in http://poignantguide.net/ruby/ close to what human speaks.

As it says ...It is coderspeak. It is the language of our thoughts.

Read the following aloud to yourself.

5.times { print "Odelay!" }

Thinking of spending some time with RoR in the coming days and may be do some hands-on during the weekend of how to use this technology in developing fairly complex web applications and leveraging the framework for a typical custom build reporting and budgeting application.

www.AppsBI.com

Stumble upon this very good blog on Oracle Applications BI modules. Nilesh Jethwa, the author of the blog covers various topics related to Oracle Applications in general and Oracle Apps BI modules in specific. There are posts on EPB, DBI and likes, for which there is hardly any material available outside Oracle own site. Author also runs a parallel blog at ITToolBox on this url. There are posts which are covering the bare tables of Oracle Applications which you should be using to extract different data like GL Balances or the tables in AR which could be used to extract sales data. On other side there are posts like this on What is ERP? which covers grave IT topics in layman terms.

Like Nilesh, I have worked for Oracle Corporation. I have been working on Oracle Business Intelligence Applications (OFA, EPB and Sales Analyzer) and other BI tools (OWB, Discoverer, Hyperion Essbase etc.) for quite sometime now. So it would be interesting to follow this blog. May be I end-up drawing some inspiration to put some good post on this space. I have put to gather a list of topics on which I intend to write something. However the list gets bigger and bigger and I hardly manage to get some time to write on some of them.

Friday, April 28, 2006

Good Oralce database blog..

Came across few postings on Oracle database at this blog. This time around author is covering on "Why Oracle Works the Way it Does". The author gives very good analogy on how Oracle internal works. Both this and this talks about Oracle table spaces/data files and Segments/extents respectively. The author helps one build quick understanding of oracle components. I was looking around for somethign similar for quite sometime. No doubt there are lot of people writting lot about how to use Oracle database effectively. This seiries in one more in the lot but a great read for wanna be DBA and Oracle database developer.

Thursday, April 20, 2006

Scientists offer 10 basic questions to test your knowledge...

Stumble upon this article which takes you through a quick test of your knowledge of Science. The questions are not too difficult and checks you awareness of common science facts and concepts. Accourding to the article the list has been prepared by a group of elite scienctists and noble laurates. I managed to get a score of 75%.

Tuesday, April 04, 2006

Few Web 2.0 things..

Stumble upon this post, which talks about what how the next generation of User Profile management by Internet services giant like Google, Yahoo, Microsoft and others would lead us to. The post talks about this original article which appeared in MSN. Seeing the pace at which thing are moving in the Internet arena, and the Web 2.0 revolution, the Internet experience can get more personalized and become intimate part of out day to day life.

Few more new things which I came to know today was digg.com and www.technorati.com. Digg is a very good Technology news, blogs and RSS aggregator. While technoratti.com is a search for the Blogosphere. Both are great webservices.

www.feedburner.com was another thing, which I stumbled upon few days back. Feedburner helps is great utility for feed management.

Every now and then I get to know about new kinds of Internet services mostly around Collaboration and Social networking. It sometime gets difficult to cope with all this.. But anyways its flow and one has to flow with it. No doubt, this new breed of services are really useful and have taken the Internet experience to the next level.

Friday, January 27, 2006

A good OWB dicussion

Came across this very good discussion on some of the bugs and missing functionalities in Oracle Warehouse Builder. The discussion though started with some of the snags in the OWB but latter on turned out to be a OWB vs Informatica comparision. There were some very good points brought out by the folks out their. I have worked with OWB extensively but have very little exposure to Informatica.
The discussion covered fundamental aspects of OWB and Informatica specifilcally and ETL in general. Worth read. Below is interesting snippet. By the way the whole discussion is worth reading

What was really of interest to me was how well the tool works in harmony with the other pieces of the data warehouse (i.e. relational sources and targets as well as external files). And for me, here is a critical point: the Informatica approach to data cleansing and movement is to act upon individual rows which are pulled from a source (often the database), scrubbed, and then placed into the target.

As the core processes of data movement and cleansing continue to be integrated into the database engine, I believe third-party vendors like Informatica will find themselves more and more on the outside looking in. Tools such as Oracle Warehouse Builder are basically frameworks around the database engine core ETL functionality and enhance user productivity by providing the glue to make all of the technological pieces work together. Is it on par with the slickest tools out there? No, not yet.... but "Paris" looms on the horizon.

Thursday, January 26, 2006

Great Yoga site

I stumble upon this very good Yoga site: www.abc-of-yoga.com. The website covers various aspects of Yoga accurately and with significant details. I learned my first lessons of Yoga through a book I got from my father titled “Be Healthy for 100 years”. The book was written in local language and was focusing more on Asanas (postures). The next set of lessons I learned was at school. We had a veteran teacher for teaching Yoga and Sanskrit. We learned some good lessons on Asanas and pranayam and were able to gracefully do the difficult asanas like Chakra asana. Since then I have been practicing some of these asanas.

Thursday, January 12, 2006

Open Web Design

While browsing through few of the entries at IT Toolbox, I came across this post talking about Open webdesign. http://www.openwebdesign.org/ has really great set of templates which could be downloaded and used for small size web applications. I think its one more noble practice happening on the Internet. Sharing information, software, advice, thoughts, ideas or to the matter anything is the mantra of the day in the Internet world.

Wednesday, January 11, 2006

Some good posts..

I started following some of the Blog categories on IT Toolbox. Apart from blogspot.com this is yet another place I regularly visit these days. Other day I stumble upon this very good entry on SEO (Search engine optimization). Though its been long I have worked with SEO related stuff, this was indeed good post putting together good list of basic SEO consideration while designing web pages. One of the notable point author makes in the beginning of article is that Search bots are enhanced these days to minimize or eliminate tricks used by web designers for a higher Page Ranking on search engines. However author goes on mentioning some things to consider while designing the web page so that its serach engine friendly and gets recognized appropriately by the search engine.

Scott Berkun, is not more guy whom I am following since an year. There are lots of discussions on Project Management, Usability Engineering etc. A series of essays Scott has written are worth read for any one from the software field. Also the PM clinic and UX clinic are good read.

Of late I also started following yet another BI community at http://www.b-eye-network.com. There are interesting set of articles and news related to BI. There are specific vertical channels like Retail, Life sciences that have good bunch of articles worth reading.

Friday, December 23, 2005

Great essay: How to be a Programmer

Few days’ back I stumble up on this very good essay titled: “How to be a Programmer: A Short, Comprehensive, and Personal Summary” by Robert L Read. The essay is must read for all those whoever are part of the “Software Programmer” tribe. It’s a comprehensive essay detailing various facets of professional life of Software engineer.

The essay deals with range of topics starting from basic debugging skills to managing projects, conceiving good design and to the extent of how to manage the people, team dynamics and personal aspects like “When to go home” etc.

The author has been very concise and focused while explaining all these topics. The best part of the article is that it’s very generic and the suggestion/recommendation holds for a person working as any role in software development with any sort of technology.

Another good part of the essay is the manner in which it is structured. All the skills required in order to be a successful programmer are grouped into three sections namely: “Beginner”, “Intermediate” and “Advanced”. The “Beginner” section talks about things like debugging, performance tuning, fundamental concepts like memory and i/o management, testing, experimenting, team skills, working with poor code, source code control etc. All are damn interesting and must read.

“Intermediate” section covers some of the soft skills, which are of paramount importance. Things like Personal skills, how to stay motivated, how to grow professionally, how to deal with non-engineer etc. and technical things like managing development time, evaluating and managing third party software, when to apply fancy computer science. Finally the “Advanced” section talks about how to make technology judgment, using embedded languages, dealing with schedule pressure, growing a system, dealing with organization chaos etc.

So all and all very interesting read.

Sunday, December 18, 2005

Outsourcing to small vendors: What to consider?

Few days back I was having mail discussion with a friend of mine on what are the prime considerations of outsourcing a project to small vendor.
I personally have worked with both a small size outsourcing firm in India named Rightway Solution (www.rigthwaysolution.com) and significantly large outsourcing partner and system integration company at Singapore.

Below are some of the snippets from our discussion:

Just wished one favor from u. I have some project which I need to outsource to india. Can you please tell me the name of the site where I can post my req, so that bidders can bid on it? Also, can u plz tell em the name of the company where u had done ur first job and u worked on outsourced projects? Also, any hints or pitfalls to be taken care of while outsourcing projects?

Regarding .. the outsourcing stuff.. You can go to some e-lacing site.
www.elance.com is the best in the league.. very sophistacted. There are some small ones too. I don't have their URL on top off my head..but you can google them.

Pitfalls.. Make sure you choose the good vendor..who has both the quality..and timely delivery.. As you know.. in the outsourting.. the Delivery model.is very critcal.. So you be in look out of the vendor which has very sound delivery model.. Won't say it shoud be very sophistacted.. but very roubust.. of good quality.

Rightway. the company I worked..has a very good track record.. They are not a big player in elance.com but they are established vendor in some of the other elancing portal... not able to recall there name.

Second aspect of the outsourcing is the communication.
-The way you communicate your requirement.
-The way you monitor the progresss of the project
-The way you gauge..the honesty and quality of their work.
-The way you get things delivered..
-The terms and condition...payment.etc.

For all this communication is the corner stone..and its of paramount importance..when you don't see the opposite party physically.

1) For a j2ee project, what delivery model can be thought of? I perceive it as a war file that can be deployed on my server. The db setup and other things would be just scripts which I need to run. All the necessary documents (like TD, FD, test plans, etc.) should be delivered as they occur in the phase of SDLC. Please suggest me if you can conceive any other model.

The final software (piece of code) and documents are as much important as much is the visibility of what is the architecture, the whole code and details about enhancing/maintaining that code. Typically the customer/client who are assigning the project to outsourcing partner are much bothered about the final software delivery and pay less attention on the documentation or the overall design of the software..
So I think the focus should be all different aspects of the software and not just final piece of code.

- Key to the delivery is the Project Plan. Depending on the size of project there should be a accurate project plan (both high level and detail level) should be agreed upon by both client and vendor. Appropriate milestone should be laid out and one should make sure there are no slippages.
- Regular status update meetings where the client and vendor sit to gather and take the stock of situation. Any issues should be raised up. Proper status report template should be followed. Close watch should be kept on the whether the milestones are met properly or not
- Regular technical discussion/ walk through specially during the design of the architecture and over all solution. Mind that you would have very less time to and rather it would be too late to rectify any design goof-ups at the latter stage of the project. So the piece of code could be just mere speck and of no use if the core is in mess.
- I would advise to have regular signoffs. This is in tandem with your Project plan. Any milestone should be signed off by the client. Vendors are typically reluctant on this since it affects there flexibility.. but I would advise you to insist on this if the size of proejct significant.

Overall one should be closely watching the overall developments happening around.. even though there is not physically proximity but through various communication channels.

I won't recommend to follow the exact SDLC cook book if you are dealing with small vendors. Yeah design documents are critical and get them done. Not so formal/sophisticated in terms of templates/formats. But they should be readable, understandable and comprehensive. It doesn't matter if you one wants to call it LLD of HLD
etc. 2 levels of design document is advisable if the size of the project is significant.


2) I am sure it would be financially beneficial to contact rightway
directly, rather than through eLance. Agreed? Also, does that approach have any cons in terms of accountability or fraud?

Yeah approaching rightway directly would be beneficial. The only advantage you get if you go via elance is that the payment is done through Elance. So you pay to elance and they pay to vendor. However it involves some commission at part of client and vendor.
Its your call. Rightway.. have pulled of from Elance since most of there business comes from other channels like references, existing clients etc.

If you are thinking of Java/J2EE, Rightway does not have much expertise in it. There forte is PHP and .Net/ASP. Just a thought here. PHP is no doubt a proven scripting language for web applications. Its robust, has lot of supports of various backends, open source and I would say its being used, grown and nurtured actively by a large community. Even a small to medium complex applications could be designed/developed perfectly and efficiently using PHP.

Rightway no doubt has tons of experience in this technology. They have best practices, a very good processes and design/development methodologoies around PHP. And PHP is easy and good to maintain. Very much
portable with wide range of web containers. So think.. rest its your call.

3) Code copyright can be enforced..right? So, a buyer of the service would own the copyright for the code and not the service provider..rt? (This should be obvious, as is the case with TCS ,Wipro, Infy, etc.)


IP (Intellectual Property) is something you can enforce on vendor. To be honest no one can stop any one using the code. However you can have some legal contract of how IP is to be managed. Usually the vendor does not sell the same code/product as it is to other client. They obviously reuse some of the generic component/ best practices in other projects and no one can stop.

M back After long time

Almost six months now. Lot happened during this 6 months. Lot of stuff to do at personal and professional front.

In July, I visited my native place, Kodinar (small town in India). Come August, I am back to Singapore. During this time we were embarking the new project:

This time it involved Hyperion Essbase. A new beast, I never worked before. The scope was to implement Sales and Marketing Analytics platform. Number of KPIs to model: 60+. 5 core dimension anchoring this 60 facts.

Most of KPIs were semi-additive. There were even some KPIs which were not additive on any of the dimension.

The client I have been working since an year now, has a very different set of metrics they want to implement in there on-going effort to create enterprise wide data warehouse.

So work involved both learning and experimenting new things. I would spend sometime in coming weeks to blog some of experiences in this last 6 months.

Tuesday, June 28, 2005

No Silver Bullet...by Frederick P. Brooks, Jr

Just came across a very good paper: No Silver Bullet: Essence and Accidents of Software Engineering. by Frederick P. Brooks, Jr.

If I am not mistaken he is the same person from IBM who has written The Mythical Man-Month.
The paper is a great read for anyone who is in software industry. Though this paper is dated back to April 1987, the principles it delineates still very much hold.

Following snippets from this paper talks about invistiblity aspect of software.

Invisibility. Software is invisible and unvisualizable. Geometric abstractions are powerful tools. The floor plan of a building helps both architect and client evaluate spaces, traffic flows, views. Contradictions and omissions become obvious. Scale drawings of mechanical parts and stick-figure models of molecules, although abstractions, serve the same purpose. A geometric reality is captured in a geometric abstraction.

The reality of software is not inherently embedded in space. Hence, it has no ready geometric representation in the way that land has maps, silicon chips have diagrams, computers have connectivity schematics. As soon as we attempt to diagram software structure, we find it to constitute not one, but several, general directed graphs superimposed one upon another. The several graphs may represent the flow of control, the flow of data, patterns of dependency, time sequence, name-space relationships. These graphs are usually not even planar, much less hierarchical. Indeed, one of the ways of establishing conceptual control over such structure is to enforce link cutting until one or more of the graphs becomes hierarchical.1

In spite of progress in restricting and simplifying the structures of software, they remain inherently unvisualizable, and thus do not permit the mind to use some of its most powerful conceptual tools. This lack not only impedes the process of design within one mind, it severely hinders communication among minds.

Frederick P. Brooks is always a great read. The way he writes on the intricacies about software industry in compare to manufacturing is too intriguing.

Saturday, June 18, 2005

Trying out OWB Java API

I have been working with OWB since a year. And the learning is still to take a back seat. During the initial days with OWB the main attraction was exploring various operators (pivot, match merge, sort, etc.), trying out all the possible things one can do from the mapping editor, configuring various objects from OWB Client, implementing things like SCD, exporting the OLAP metadata for use with BI Beans, process flows etc.

Latter on, the focus shifted to Runtime and Design time browser. Both the browsers were significant features of the OWB. RAB (Runtime Audit Browser) is the great place to see the audit logs, the error messages etc. etc. Design time has two great piece (Impact diagram and Linage diagram), which could be of great help while analyzing the impact of any change in the OWB mappings.

Then came the exposure to the various scripts under /owb/rtp/sql. The scripts were of great help understanding runtime platform and how to manage the RTP service. Apart from that the scripts like sql_exec_template.sql, abort_exec_request.sql and list_requests.sql were really handy in managing the individual runs of mappings.

Then came the time to explore various tables/views in the design time and runtime repositories. Later on I managed to get a chance trying my hands on OMBPlus. And man, it was just too cool working with OMBPlus. OWB Client (gui) is extremely unusable if there is some repetitive task to be done. For example you want to substr all the attributes to 30 characters. And assume that there are 50 such attributes which needs to be sub-string. What does one do? Drop an expression operator. Get all this attributes into the input group. Then one by one add 50 attributes in output group. Then change the expression property of all this 50 attributes to do the substr() of the corresponding input attributes. Now this is heck of work. At OMB side its just one small script which you got to write to do the whole thing. So if there is anything repetitive and tedious OMBPlus is the answer. Even taking backup of the repositories in to MDL is repetitive. One can write a small bat job using OMBPlus to take care of it. Apart from this, there could be many more things which OMBPlus accomplishes for you like creating template mapping, deploying the object, synchronizing the objects, importing the metadata etc.

So, over the period I kept learning all the different pieces of OWB, which I feel is an extensive suite and provides a comprehensive capabilities for any ETL and data integration task. And still I had few more left out. One of this was Java API to manipulate the metadata. Oracle exposed a public Java API to manipulate OWB design and runtime metadata, deploying and running the mapping, etc. Apart from all the capabilities of OMBPlus, Java API has some additional capabilities. In reality OMBPlus in turn uses Java API to do the various manipulation to the metadata.

So today was the day to get my hands dirty with Java API which came along with OWB 10g1 and onwards. The first thing I did was to find out some document and all I was able to manage was http://download-west.oracle.com/docs/html/B12155_01/index.html, the Javadoc for the API. Typically the Javadoc just has the description of all the classes, interfaces, methods etc. The doc is not a step-by-step tutorial on how to write a simple Java program using the OWB Java API.

The API is very big and so is the doc. The above link leads you to a list of 25 odd java packages to do various things. However there is no hint on where to start.

I brought up my Eclipse (www.eclipse.org) workbench and created a simple Java project named OWBApi. There was a bit struggle in locating where the jar file for the Public API resides. I managed to locate it under /owb/lib/int/publicapi.jar. I added this jar to the build path of the OWBApi project by going to menu Project->Properties -> Java Build Path -> Libraries and Add External Jars.

Back to the Java doc for the OWB Java API:
oracle.owb.connection was the first package I hit. Everything has to start with connection first. After bit of scramble here and there I managed to put together following piece of code in ConnectOWB.java


import oracle.owb.connection.OWBConnection;
import oracle.owb.connection.RepositoryManager;

public class ConnectOWB {
public static void main(String[] args) {
try{
oracle.owb.connection.RepositoryManager rm=oracle.owb.connection.RepositoryManager.getInstance();
OWBConnection owbconn = rm.openConnection("dtrep",
"dtrep","localhost:1521:orcl",rm.MULTIPLE_USER_MODE);

if (owbconn != null ){
System.out.println("Connection Establishied..");
}

}
catch (Exception e)
{
e.printStackTrace();
}
}
}


When running the above, I got following error message:

java.lang.NoClassDefFoundError: oracle/wh/repos/impl/foundation/CMPException
at ConnectOWB.main(ConnectOWB.java:23)
Exception in thread "main"

This means the build path is missing some jar which contains oracle/wh/repos/impl/foundation/CMPException. Now how to find this. Tried searching for this error on OWB forum at OTN, etc. but of no help. Finally an idea clicked to add all the jar under /owb/lib/int. Doing so I was able to get rid of the NoClassFound error message but ended up with one more error message saying that
unable to located Compatibility.properties file.

Where to go now? I tried to search this file under OWB home and was able to locate it under /owb/bin/admin. Inorder to make this file avilble to my java program I added one more entry in the build path for this directory. Adding a folder to the build path is as good as adding a Jar but instead of selecting the Add External Jar one has to click Add Class folder. Spicify the directory: /owb/bin/admin.

That’s it. I tried compling and running the program again. And it’s through. I was able to establish the connection to design time rep.

After oracle.owb.connection, the next hit was oracle.owb.project. I manged to do some more things like getting the list of project, creating a project, setting the active project etc. Following program displays the list of Projects in the design time repository.

import oracle.owb.connection.OWBConnection;
import oracle.owb.connection.RepositoryManager;
import oracle.owb.project.*
;

public class ConnectOWB {
public static void main(String[] args) {

try{
oracle.owb.connection.RepositoryManager rm=oracle.owb.connection.RepositoryManager.getInstance();
OWBConnection owbconn = rm.openConnection("dtrep",
"dtrep","localhost:1521:orcl",rm.MULTIPLE_USER_MODE);

if (owbconn != null ){
System.out.println("Connection Establishied..");
}

ProjectManager pmgr = ProjectManager.getInstance();
String[] projlist=pmgr.getProjectNames();
for(int i = 0 ;i<projlist.length;++i)
{
System.out.println(projlist[i]);
}

}
catch (Exception e)
{
e.printStackTrace();
}
}
}


And the list goes on. Java API seems to be a good option to write the striped down interface like OWB for some special set of users who really don’t need the whole OWB client to

Java language is a proven language for writing the GUI application. It has a rich set of libraries for accessing developing user interfaces, network programming, database programming.

One can use this API and end up writing a browser-based interface to manipulate the OWB metadata. Or may be even creating new mappings with some specific templates and stuff. Or atleast writing an interface to run a job, deploy a mapping, change some metadata etc. This really reminds of a new feature called Expert, which is coming along with OWB Paris.


Yeah, so getting back to my learnings in OWB I have some more in list. Next would be exploring the Appendix:C of the user guide. The appendix talks about extracting the data from XML data sources. Next is to use the AQ (Advance Queues) and build up the understaning of pulling the data from Appilcations (SAP). So still long way to go.


Thursday, June 16, 2005

How to get started with the new technology/tool etc.

Learning new tools and technologies has become part of daily chores of any IT professional. There is no way out. Or there is no good reason of why one should not learn new things. I personally am a tech savvy guy and always in lookout of learning new things. The interest is not just to learn things pertaining to data warehouse and BI but everything, which comes on way. The only thing that the new tools/technology I learn should have some fundamentals or concepts to take home.

During this last 6-7 years of being into IT, I have learned numerous theories, technologies, programming languages, tools etc. Most of them were through self-learning. But this self-learning was dependent on all my previous learnings, which I inculcated in the past and without which all this self learning would not have been possible. Today I just picked one more tool/technology to build some understanding on it. I don’t have the access to the software but just the documentation. This is one among few tools/technology I am trying to learn for which I don’t have the access of the software. Though I have hands on extensive hands-on experience on a similar kind of technology by another vendor.

This whole thing lead me to think of how can one approach taking up new tool/technology. Possibly three ways which came into my mind:

1. First hit the document.. get some background. and then come to the tool/hands-on and then again go back to the manuals/references. .. an then back to hands on.. May be over the period doing both things simultaneously.
2. First hit the tool ..let your intuition take over the wheel first.. play around stretch your understanding/intuition... and then come back to references/manual/docs/some text and then back to the tool. Over the period both doing both things simultaneously.
3. First attend some seminar ,some talk, some discussion ( as good as 1 but instead of text you are get into more live things) and then hit the tool may be then back to the manuals tools.. come back to tool/hands-on then go back to discussion and so forth. May be I call it Spaghetti approach. In this approach it could be that you start with Books first and then tools an then talks or any combination.

Which to choose?? Time and availability of resources can give the right call for this.. I keep trying all this approaches. Most of the times approach 2 is a good deal for me. Approach 1 is something we have been trying since the college days. First read about the “c” language, listen some lectures... and then get to the labs for some hands-on. And that was good since one didn't had so many fundamentals/concepts built up, not so much of exposure to the tools/languages of similar kind. Again like all my postings, there is no need to reach to conclusion of which is better and which not. Depends like everything else. My idea here is just to bring out some points.

Monday, June 06, 2005

Funky Business

During this weekend I got hold of this book Funky Business and man I can't keep away myself from getting it done. Now I don't want to end up writing one more review of this book, buts just thought of sharing some of the intriguing things I liked about this book. The authors have really brought in lot of wisdom of how business should be run in the 21st century. And all that in different style of writing. Everything is funky about the book, the examples, the style of writing, the wisdom, the content, the authors. Its just cool piece.

There were lot of striking things, lot of striking concepts which just hits your nerves like a sharp tool, lot of striking examples (the one defining niche market was: group of lawyers who are interested in pigeon races). And no doubt the book has tons of facts (GM tried producing car stereo and that didn’t worked out for them, or a dentist slur company which has 50 worlds market share is just run by 85 folks). Now all that is interesting. And above all lot of learning one can draw out from this. The one good thing about this book (or may be bad for someone’s) is that its very concise and says 10 things in 5 sentences. So one has to just keep reading it again and again to appreciate all what it has to offer.

Quick Linux Recipe

The other day I had a task to install some avtaar of Linux on an WinXP machine. One of my friend wanted to get started with Linux. He wanted to do some hands-on running, various commands, get hold of some basics of how Linux works, and gradually some further details like file system, Linux daemons, networking stuff in Linux and so forth.

Need was to install Linux on top of host OS WinXP, so that he can keep working on XP and switch to Linux for doing some hands on, etc. etc. The desktop he was running was bit out of time. 128 MB ram and 500 mhz of CPU. And on top of this we have a mammoth WinXP running.

VMWare was there to create a virtual machine on top of which the I was to install some distribution of Linux. Redhat was big bloat (4 disc for FEDORA) plus lot of space to set up whole thing, plus it will be killer to the CPU. So the idea was to get hold of some mini Linux distribution, which does not take whole lot of space and can get installed quickly.

I came across a roster of mini Linux distribution. But there was this BeatrIX which clicked to me. BeatrIX was cool piece of Linux bundling (<200 MB) with no need to setup since it runs directly boots from the CD. It has all the pieces which one would need to get started with Linux: Gnome, text editor, browser, terminal etc. etc.)

I downloaded the ISO image of BeatrIX. Instead of burning it to CD, I set of my VMWare CDROM to read this iso image file. And that’s it. The whole thing took less the 30 minutes. Just to recap the whole recipe:

1. Download VMWare Workstation 5 for Windows. Install it. Register and get the evaluation license key from the VMWare site (mind that its just for 30 days).
2. Download BeatrIX to some location under your file system.
3. Launch VMWare. Create a virtual machine as “Other Linux Distribution Kernel2.6.. “.
4. Modify the CD Rom device in VMWare to read from the ISO image and specify the file location to the downloaded copy of BeatrIX.
5. That’s it. Click Start. You have your Linux set up. The BeatrIX Linux OS does not need any installation since its boots up directly from CD.

By the way, BeatrIX seems to be have built and inspired by interesting set of people and cats. Check this out there site.

Sunday, June 05, 2005

What is xml?

So this 3 letter word is doing storms in the IT world since its inception in late 90's. Now what is it all about? I hear lot of folks talking around, defining, trying to understand, trying to explain other, of what XML is. Even I my self have indulged in all such discussions. I kept hearing lot of definitions floating in the air some saying "Its the standard to encode data", "its enhanced version of html", "its extensible HTML, you can create your own tag", but why on the earth would I need to create these tags? What for?

The understanding which I build up in this due course of discussion and reading was that "XML is a standard way to encode the data which is pertaining to anything ranging from transaction details, list of entity, a message for some application, configurations, metadata etc. and the only way it differs from the a simple text file is that in XML data is stored in hierarchical fashion and an XML document is bound to some schema or DTD which specifies the structure and content of this hierarchy." XML is a way to package a data. Now this packaged data could sit in a file or network packet or a message or a database table or anything.

I didn’t worked with XML per se. As such there is nothing like working with XML. XML is not a programming language which one can use to create some application or neither its meant for presentation like HTML. As some one has said that one will encounter XML everywhere. Even when your car will break down, it will send an message in XML to the nearest service center for necessary help.

XML is meant for nothing in specific but for everything. I kept seeing XML everywhere in the last couple of years,
- Configuration of various applications/servers
- Web services are sending request and response in XML
- The report I create using some tool gets stored in XML file. Its not just reports but any meta data generated using any wiggy wizard tools gets stored in the XML.
- I write my descriptor file of EJB in xml, my strut config is in XML
- The WML is again an XML
- The process flows are getting stored in the XML
- The presentation information is getting stored in the XML and is transformed to particular rendering device using some translation
- I export data from the database in XML and import it to any database.

These and many more. I wonder why use XML everywhere if the bare simple text files can do the same? Okay what could be a bare text file look like which stores the list of books and their details:

Option 1 (attribute value):

Book: Abc
Author: Xyz1
Price: 100
Pages: 252
Book: Abc1
Author: Xyz2
Price: 150
Pages: 531

Option 2 (comma separated):

Abc, Xyz1, 100, 252
Abc1, Xyz2, 150,531



And may be there could be some more.
For both of the above options the application has to make necessary assumption when consuming or generating this text format. In the first one, all the attributes for one book should be placed together vertically and in the second all the attribute for a book should be place horizontally together in a particular order. And in case if this file needs to be extended to store some more attributes for a book, for example Publisher information. So what all needs to be changed? Application? File? We understand it. It will be heck of work.

On the other hand, XML is also the bare text with some structure and some syntax to follow. That’s it. Something, which I have found till now, which is bit convincing to me and which stands XML better then simple text encoding:
1. The structure of a xml document is extensible without effecting much of application. You can extend the xml document to store some more information without effecting the application which is using it
2. The data is stored in hierarchical fashion something like:
<?xml version="1.0"encoding="utf-8"?>
<Books xmlns="http://tempuri.org/XMLFile1.xsd">
<Book>
<Name>Abc</Name>
<Author>Xyz1</Author>
<Price>100</Price>
<Pages>252</Pages>
</Book>
<Book>
<Name>Abc1</Name>
<Author>Xyz2</Author>
<Price>150</Price>
<Pages>531</Pages>
</Book>
</Books>

This could be very well extended to store the new attributes without really bothering the application.

3. Availability of lot of parsers and DOM (Document Object Model, API for accessing for processing XML document) for various programming languages. So generating and consuming XML document is easy

There are two guidelines, which every XML document has to follow:
1. The XML document should be correct. This means every opening tag should have a closing tag. The structure should be correct. And the tags are case sensitive.
2. The XML document should be valid. This means that the arrangement of tags, there attributes and there values have to follow certain scheme. This scheme is specified in the DTD or XML schema, which is associated with the XML document. In the above example it is http://tempuri.org/XMLFile1.xsd which specifies the schema of the XML document.

One can Google and find tons of commentary on XML, XML toturial, applications of XML, current happenings etc. www.w3c.org is the place to get the latest on what’s happenings in XML world.

Saturday, June 04, 2005

Handling of testing and QA artifacts at the source system while building ETL

Now this is interesting. Your source system (the production database) has lot of test data spread across the various entities. A typical online portal can have routine set of test cases run every day on the production system to check the consistency of system. So how does one handle all this test artifacts while building the ETL? Should this be treated as part of data cleansing? May be it should be or may be it needs more serious attention then just cleaning them away.

Handling test data in the ETL involves two aspects: one to identify and segregate the test data and second to track the test data at a prescribed location. This location could be some separate set of tables in data warehouse it self (if its an requirement) or some log/audit files or even the tables which store the regular data with a tag saying that they are test supplier/buyer/product and not the actual.

Segregating test data from the production data depends on flagging done at the source side or some convention followed while generating the id for the test data. For example id starting with 65xxxxxx is always test data. Another way of segregating test data would be a lookup table residing in the source system or staging which contains the list suppliers/buyers/products etc. which are test data and corresponding transactions are test transactions.

If the case is just to identify the test data and filter it before bringing into the staging or DW, life would be perhaps easy. However if the need is to bring the data in data warehouse with some identification to separate it out, there could be two possible ways to do it: populate the data in separate set of tables or in the same set of table with some tagging. The latter has an advantage because it saves the extra ETL at the cost of the one extra flag. The second approach also aligns the test data with the regular data hence the same constraints checking and data capturing, ETL could be used. But there could be tough times handling the test data with second approach if it does not follow the prescribed application/business logic. This could be due to some data patching done from behind to run through some test cases or any special provision in the application logic. First approach stands out to be better for this case. There should be enough balance maintained such that main ETLs populating the regular data does not get complicated just because its handling lot of exception for test data. The best deal here would be do segregate the test data at the first place and put it in the separate table which in turn could be used as lookup for regular ETL.

There is no definite thumb rule,(as such there are no thumb rules) of handling the test data in the source system. All depends on the nature of test data and identifying it and the way it needs to be tracked.