Good Smaritan

May the best things in life happen to Mr. Manjunath. The guy who just dropped by my office and gave me my “high-value” cheque. You can still find good guys in this world.

Manju, says that he found the cheque on the staircase of the ATM booth. I had clearly put the cheque in the box. So the theory of the cheque falling off from the clearing team proves right. What a negligence on their part ! If Manjunath hadn’t given it to me, then I would be wondering for nearly a month as to what happened to that cheque. Even HDFC would back off says “hey we didn’t find your cheque”. Thanks again to Mr Manjunath for saving the day.

One lesson learnt – don’t put the cheques in the drop boxes. The good old way of depositing it over the counter is what I’ll do from now on.

Cheque Drop-Box

Just got a call from a person who has my high value cheque with him !

Just yesterday, I had put the cheque in the “high value drop-box” on M.G.Road branch of HDFC. I was shocked for a moment and didn’t know what to answer. The person says that he got the cheque today morning near the ATM of the M.G.Road branch. The only thing probable is that it got dorpped off the box during the clearing time. Thank God it’s a/c bearer cheque.

I’m just hoping that the person who called me will come near my office by 6:30 p.m. and hand over the cheque. Crossing my fingers that he really turns up !

Bhentures

It’s an event organized by Ventures in Bangalore for corporate. Oracle India has been taking part in this for the past 3 years (if my memory serves me right). This year it was represented by a handful of people. Otherwise, it was just 2 members; Rishi and Bobby. Both from ISC and very good players of shuttle. Bobby has some “slip disc” kind of a problem and his lower back is always under stress. Still he manages to take part in this event and be competitive at it. Any other person would have just given up due to the injury, but Bob goes on and on…

About Rishi, he is simply superb. An Asian Games gold medalist. So you can imagine how competitive his game would be. Together they thrashed every other team when oracle first made its debut in the event. After the initial year, it was struggle for these guys. There are so many entries for a company; singles, team event, doubles, mixed doubles and so on. These 2 can’t just enter all of them and play it. As the event is organized for a couple of weekends and they’ll be having back-to-back matches. Still these two just keep collecting prizes for the entries they make.

This year, it was different. They have found another doubles partner who is known for his “smashes”. Rishi says, that he needs a partner who can just go on smashing. The new find just does that. When I watched the game yesterday, it was nice to see them going. With my poor knowledge of shuttle, I think they would need some more matches to be at their best. The good thing is that they both appreciate each other’s game, which is a very good sign in any partnership. Hopefully they’ll just keep getting better and better to get home the championship.

One-on-One with Tech Lead

Had an interesting discussion with tech lead of a project out here. I was asked to setup a local 10gRAC Cluster for their use. After which I didn’t have a roadmap of my involvement with this project. When we sat down to chalk one, the first question posted to me was “are you interested to do these?” – for which I had no answer (. Kept wondering where did this question come from?

The previous project was a total disaster. There were lots of issues from the DB. The final result; application didn’t scale and performed very poorly in terms of performance. So the tech lead wasn’t happy with the DBA help on that and hence the question was fired to me (

Well, I’ve never been involved so closely with an application design. Also I’ve been reading Tom Kyte’s book on design which just makes one point “Design for performance” and “DON’T TUNE FOR PERFORMANCE”. So this seems to be a perfect opportunity to follow the “best practices”.  I’m all eager to start off on this one, which means that lots and lots of reading, testing and documenting stuff…

One Hour Closure

If you are working in support you would know exactly what the title meant for your appraisal 🙂

How would it be that the support analyst says that “I haven’t done a 10.1.0.4 patchset installation”…man that was mean. I had to convince him to use owc to check it out for himself. Imagine a customer pestering for the usage of owc rather than the analyst.

After checking out on my own desktop, the analyst had to say “this is the first time I’m seeing this kind of an issue” . I was very glad to hear that. But I definitely need to appreciate one thing from him. In my urgency I had overlooked that the root10104.sh file was present on the second node. Surprising !! He did catch that and it helped me. I really thank him for that. Didn’t wait for the guy to say anything more, scp the file to node1 along with another $ORA_CRS_HOME/install/patch10104 dir and finished my patchset application.

About the SR, I gave him the option of a ‘1 hr closure’…let everybody be happy. I got my job done and let me get something for that. Couple of things happen here and there, but it always feels nice there is a ‘win-win’ situation. With that we rest for the day.

First iTAR

Today I opened my first iTAR (known as SR these days) for a missing root10104.sh file. Got my 10g RAC installed on a couple of nodes, then created a small test db. Thought I’ll check out 10.1.0.4 on it. Did all the things judiciously(reading the readme from top to bottom, following every instruction). After the installation finished, as usual I went in search of this root10104.sh and found that to be not present where it was supposed to. Did a couple of checks here and there, the guy just doesn’t exist there.

IMed a couple of ORCL guys and they were aware of it. So I thought, ok let me go formal with it and opened my first iTAR.

I’ve been there for some time and I can easily recognize when an analyst is buying time. When you get an immediate response for an RDA ouput, thats one of them. Initially looking at the update, I was really pissed ! But then there are cases when an RDA really helps an analyst to get the work done. So I gave him the RDA output and haven’t heard from him since then !

Actually I don’t expect him to get back to me so early. When we were starting our days as support analysts, we ensured that the TAR gets updated with the right kind of information. The analyst is not a super human being. He is just another guy – happens to be equipped at handling these kind of issues. But still one needs to remember that product support is given out as a ‘service’. And in any service oriented, keeping the customer cool about things is very important. All it takes is a line which can say ‘Hey I haven’t got this error…let me lookup and keep you posted on it’.

As I was writing this…the guy came back with a response which said “Some of the files haven’t been copied to the second node during installation..can you send in those logs?”

Man Give me a BREAK !!!!

Production Disasters

Just finished a marathon weekend. On thursday, one of our production server’s disk decided that it was time and went foooo ! The box has a hot-swap thing, but something happened(which is always the case in production) that even the hot-swap was not working. Since the storage was from the filer, all the production (8 in number) databases were down. The sys admins got the h/w fixed and it was time for oracle databases.

Since I was very new to the organization, I wasn’t allowed to check what was the extent of damage to oracle databases. When the db was started, dbwr complained that couple of datafiles are missing. So they got it from the recent backup and tried a recovery.

Should have been pretty simple considering that it was just another recovery of datafiles. The problem was with the way recovery was attempted on them. After restoring the datafiles, one doesn’t attemp an incomplete recovery on them. Scanning the alert.log of the db in question, I found that the statement used was ‘recover database until cancel’.

How would you expect the db to recover from this. The rest of the datafiles are at time say t1 and these restored datafiles need to be in t1-x. Just a tiny mistake is enough to get your weekend schedule go haywire.

What would it take to have a tested procedure for disasters. Something which will give the on-call dba a checklist to analyze the extent of damage (if any) and what are the steps that are required to get the db back in production?

After this incident, I’ve made a personal commitment to have a ‘Recovery Plan’ which would address most of the crashes and the steps that are need to get the db up. This would definitely make the life of the on-call dba a lot more helpful.

The Bengalûru Oracle Meetup Group

Oracle Meetups

Before I got pulled into our production db recovery, I was pretty much idle and was waiting for root_squash thing to be resolved. Just did a search for any local oracle user’s group and found this The Bengalûru Oracle Meetup Group. Currently there are about 13 members in this meetup group. But the group lacks an oraganizer as of now. So that..what the heck, lets be one and learn to co-ordinate some stuff…an oraganizer needs to pay $19/Month for it.

I’m thinking if I can get it going with a couple of other friends from oracle, why not pay so much and have a group and lets see how that goes. I’m enquiring it with my friends and once I get something concrete, will go ahead with it. I can get my friends to deliver presentations, present papers, give tips. I’m sure all that would be useful for the others.

To make a first promotion to The Bengalûru Oracle Meetup Group I’ve included their logo here 🙂

web site hit counter

root_squash + 10g RAC

If you were like me (15 mins ago), then you would be wondering what the hell is root_squash and what relation does it have with 10g CRS installation. Its due to this, that I’m not able to finish my 10g CRS install.

root_squash happens to be a NFS mount option. Currently I’m setting up a 10g RAC cluster locally in Bangalore, so that the developers have a first hand meet with it. The infrastructure consists of using a NFS storage from NetApp filer, 2 nodes in the cluster. After getting the basic configuration ready, I started with CRS install. It went fine all the way till root.sh execution. When root.sh was run from the first node, I got the following output:

[root@romerac1 crs_r1]# ./root.sh
Checking to see if Oracle CRS stack is already up…/etc/oracle does not exist. Creating it now./bin/chown: changing ownership of `/ocr/voting_ocr/ocr.file’: Operation not permitted

Being in support for some time and doing multiple installations, I was surprised to get an err on chown and that too for the root user ! During the debugging session, I noticed that a file created by root has the nfsnobody for the user & group. Whereas a file created by oracle user has oracle:dba. Why aren’t the files created by root user being owned by root ?

Thats when the security part of NFS comes into play. With root_squash enabled, root users don’t become root user on the filer mount point…for obvious security reasons. So any change of ownership to root user are not permitted and hence the error message.

As of now I’ve raised a priority ticket with the storage ops to get that changed…nahh…no security concerns here as we are way behind many many firewalls (hoping they stay put).

With my first install of 10g RAC in Yahoo! I learnt a new thing about nfs moutpoint. In Oracle all these would have been taken care. The plus point is I get to see ‘new errors’ 🙂 which is good for learning. So hoping for a very exciting time at Y!.