Thursday, June 07, 2007

lucene in php & lucene in java

Something i found out while solving some issue from Mr. Nguyen from vietnam.

He used lucene-php in zend framework for building a lucene index and searching on the index, and was facing issues with search times. It turned out that mysql full text index was performing better than lucene index.

So i did a quick benchmark and found the following stuff

1. Indexing using php-lucene takes a huge amount of time as compared to java-lucene. I indexed 30000 records and the time it took was 1673 seconds. Optimization time was 210 seconds. Total time for index creation was 1883 seconds. Which is hell lot of time.

2. Index created using php-lucene is compatible to java-lucene. So index created by php-lucene can be read by java-lucene and vice versa.

3. Search in php-lucene is very slow as compared to java-lucene. The time for 100 searches are -

jayant@jayantbox:~/myprogs/java$ java searcher
Total : 30000 docs
t2-t1 : 231 milliseconds

jayant@jayantbox:~/myprogs/php$ php -q searcher.php
Total 30000 docs
total time : 15 seconds


So i thought that maybe php would be retrieving the documents upfront. And changed the code to extract all documents in php and java. Still the time for 100 searches were -

jayant@jayantbox:~/myprogs/java$ java searcher
Total : 30000 docs
t2-t1 : 2128 milliseconds

jayant@jayantbox:~/myprogs/php$ php -q searcher.php
Total 30000 docs
total time : 63 seconds


The code for php search for lucene index is:

/*
* searcher.php
* On 2007-06-06
* By jayant
*
*/

include("Zend/Search/Lucene.php");

$index = new Zend_Search_Lucene("/tmp/myindex");
echo "Total ".$index->numDocs()." docs\n";
$query = "java";
$s = time();
for($i=0; $i<100; $i++)
{
$hits = $index->find($query);
// retrieve all documents. Comment this code if you dont want to retrieve documents
foreach($hits as $hit)
$doc = $hit->getDocument();

}
$total = time()-$s;
echo "total time : $total s";
?>



And the code for java search of lucene index is

/*
* searcher.java
* On 2007-06-06
* By jayant
*
*/

import org.apache.lucene.search.*;
import org.apache.lucene.queryParser.*;
import org.apache.lucene.analysis.*;
import org.apache.lucene.analysis.standard.*;
import org.apache.lucene.document.*;

public class searcher {

public static void main (String args[]) throws Exception
{
IndexSearcher s = new IndexSearcher("/tmp/myindex");
System.out.println("Total : "+s.maxDoc()+" docs");
QueryParser q = new QueryParser("content",new StandardAnalyzer());
Query qry = q.parse("java");

long t1 = System.currentTimeMillis();
for(int x=0; x< 100; x++)
{
Hits h = s.search(qry);
// retrieve all documents. Comment this code if you dont want to retrieve documents
for(int y=0; y< h.length(); y++)
{
Document d = h.doc(y);
}

}
long t2 = System.currentTimeMillis();
System.out.println("t2-t1 : "+(t2-t1)+" ms");
}
}


Hope i havent missed anything here.

Wednesday, May 23, 2007

installing zte mc315 on linux

Here are 4 steps to get the reliance aircard ZTE MC315 on linux. I got it running on ubuntu 7.04.

1. Insert the card and look at dmesg
output : 0.0: ttyS3 at I/O 0x2e8 (irq = 3) is a 16C950/954

This says that your card was detected properly.
You can also run "pccardctl info"
output =
PRODID_1="CDMA1X"
PRODID_2="CARD"
PRODID_3=""
PRODID_4=""
MANFID=0279,950b
FUNCID=2
PRODID_1=""
PRODID_2=""
PRODID_3=""
PRODID_4=""
MANFID=0000,0000
FUNCID=255


2. edit /etc/wvdial.conf

[Dialer Defaults]
Modem = /dev/ttyS3
Baud = 57600
SetVolume = 0
Dial-AT-OK ATDT Command =
Init1 = ATZ
FlowControl = Hardware (CRTSCTS)
Phone = #777
Username =
Password =
New PPPD = yes
Carrier Check = no
Stupid Mode = 1

3. Run set serial
(download and install setserial - if you dont have setserial)
"setserial /dev/ttyS3 baud_base 230400"

4. run wvdial
wvdial

So to get it running everytime you will need to create a shell script doing steps 3 & 4. Say mydial

---------------
setserial /dev/ttyS3 baud_base 230400
wvdial
---------------

PS
This method is also not perfect. I got one card working with this procedure, but could not get another card working using the same procedure.

Let me know if this helps you getting your ZTE MC 315 card running on linux.

PS 2
And to add to it, finally i have got my card running.
The trick is simple - you have to play with the UART and the baud_base.

Just run setserial on your machine and it will give the following output for UART and baud_base

* uart set UART type (none, 8250, 16450, 16550, 16550A,
16650, 16650V2, 16750, 16850, 16950, 16954)
* baud_base set base baud rate (CLOCK_FREQ / 16)


What i did was that i kept on changing my UART and kept the baud_base as 230400. I changed my UART to 16850 and then to 16750. And my card responded at 16750. And i got connected to the internet

So, My dialup script does this before doing wvdial

setserial /dev/ttyS3 uart 16750
setserial /dev/ttyS3 baud_base 230400


And this gets my card running. So to get your card running, i would suggest you try setting different UART and baud_base...

Anger Management

Having a bad day?

When you occasionally have a really bad day, and you just need to take It out
on someone, don't take it out on someone you know -- take it out on someone
you don't know. I was sitting at my desk when I remembered a phone call I had
forgotten to make. I found the number and dialed it.

A man answered, saying, "Hello."

I politely said, "Could I please speak with Robin Carter?"

Suddenly, the phone was slammed down on me. I couldn't believe that Anyone
could be so rude. I realized I had called the wrong number. I tracked down
Robin's correct number and called her. I had accidentally transposed the last
two digits of her phone number. After hanging up with her, I decided to call
the 'wrong' number again.

When the same guy answered the phone, I yelled, "You're an *******!" and hung
up.

I wrote his number down with the word '*******' next to it, and put it in my
desk drawer.

Every couple of weeks, when I was paying bills or had a really bad day, I'd
call him up and yell, "You're an *******!" It always cheered me up.


When Caller ID came to our area, I thought my therapeutic '*******' calling
would have to stop. So, I called his number and said, "Hi, this is John Smith
from the Telephone Company. I'm just calling to see if you're familiar with
the Caller ID program?"

He yelled, "NO!" and slammed the phone down.

I quickly called him back and said, "That's because you're an *******!"

One day I was at the store, getting ready to pull into a parking spot.
Some guy in a black BMW cut me off and pulled into the spot I had patiently
waited for. I hit the horn and yelled that I had been waiting for that spot.
The idiot ignored me. I noticed a "For Sale" sign in his car window . . so, I
wrote down his number.

A couple of days later, right after calling the first ******* ( I had his
number on speed dial), I thought I had better call the BMW *******, too.

I said, "Is this the man with the black BMW for sale?"

"Yes, it is."


"Can you tell me where I can see it?"


"Yes, I live at 1802 West 34th Street . It's a yellow house, and the car's
parked right out in front."

"What's your name?" I asked.

"My name is Don Hansen," he said.

"When's a good time to catch you, Don?"

"I'm home every evening after five."

"Listen, Don, can I tell you something?"

"Yes?"

"Don, you're an *******."

Then I hung up, and added his number to my speed dial, too. Now, when I had a
problem, I had two assholes to call.
But after several months of calling them, it wasn't as enjoyable as it used to
be. So, I came up with an idea. I called ******* #1.

"Hello."

"You're an *******!" (But I didn't hang up.)

"Are you still there?" he asked.

"Yeah," I said.

"Stop calling me," he screamed.

"Make me," I said.

"Who are you?" he asked.

"My name is Don Hansen."

"Yeah? Where do you live?"

"*******, I live at 1802 West 34th Street , a yellow house, with my black
Beamer parked in front."

He said, "I'm coming over right now, Don. And you had better start saying your
prayers."

I said, "Yeah, like I'm really scared, *******."

Then I called ******* #2.

"Hello?" he said.

"Hello, *******," I said.

He yelled, "If I ever find out who you are...!"

"You'll what?" I said.

"I'll kick your ass," he exclaimed.

I answered, "Well, *******, here's your chance. I'm coming over right now."

Then I hung up and immediately called the police, saying that I lived at 1802
West 34th Street, and that I was on my way over there to kill my gay lover.

Then I called Channel 13 News about the gang war going down on West
34th Street.

I quickly got into my car and headed over to 34th street .

When I got there, I saw two assholes beating the crap out of each other
in front of six squad cars, a police helicopter, and the channel 13
news crew.

NOW, I feel better - This is "Anger Management" at its very best.

Friday, May 11, 2007

mysql multi-master replication - Act II

SETUP PROCEDURE
Created 4 instances of mysql on 2 machines running on different ports. Lets call these instances A, B, C and D. So A & C are on one machine and B & D are on another machine. B is slave of A, C is slave of B, D is slave of C and A is slave of D.

(origin) A --> B --> C --> D --> A (origin)

Each instance has its own serverid. A query originating from any machine will travel through the loop replicating data on all the machines are return back to the origniating server - which in effect will identify that the query had originated from here and would not execute/replicate the query further.

To handle auto_increment field, two variables are defined in the configuraion of the server.

1. auto_increment_increment : controls the increment between successive AUTO_INCREMENT values.

2. auto_increment_offset : determines the starting point of AUTO_INCREMENT columns.

Using these two variables, the auto generated values of auto_increment columns can be controlled.

Once the replication loop is up and running, any data inserted in any node would automatically propagate to all the other nodes.

ISSUES WITH MULTI-MASTER REPLICATION & AUTOMATIC FAILOVER

1. LATENCY

There is a definite latency between the first node and the last node. The replication steps lead to the slave reading the master's binary log through the network and writing it to its own relay-bin log. It then executes the query and then writes it to its own binary log. So each node takes a small amount of time to execute and then propagate the query further.

For a 4 node multi-master replication when i replicated a single insert query containing one int and one varchar field, the time taken for data to reach the last node is 5 ms.
This latency would ofcourse depend on the following factors
a) amount of data to be relicated. Latency increases with increase in data
b) Network speed between the nodes. If the data is to replicated over internet, it will take more time as compared to that on nodes in LAN
c) Amount of indexes on the tables being replicated. As the number of indexes increase, the time taken to insert data in tables also increases. Increasig the latency.
d) The hardware of the machine and the load on the machine will also determine how fast the replication between master and slave will take place.

2. AUTOMATIC FAILOVER

Automatic failover can be accomplished using events and federated tables in mysql 5.1. Federated tables are created between master and slave and are used to check connection between master and slave. An event is used to trigger a query on the federated table which checks the connection between master and slave. If the query fails, then a stored procedure can be created which should chunk the master out of the replication loop.

So suppose in the loop shown above, A goes out, then the event on B would detect that A is out and would make D as a master of B.

This is the theoretical implementation of automatic failover. Practically there are a few issues with its implementation.

a) for failover, you need to know the position of master's master from where the slave should take over. (So B should know the position on D from where it has to start replication). One of the ways to do this is that each slave logs its master's position on a table which is replicated throughout the loop. But this again is not possible using events & stored procedures - cause there is no query which can capture the information available from "SHOW SLAVE STATUS" query in a variable and write it in a table. Another way to do this is to have an external script running at an interval of say 10 seconds which logs this information and checks the federated tables for any disruption in the master-slave connection. With this methodology, the problem is that there could be a 10 second window period during which data can be lost.

b) You could also land into an infinite replication situation. Lets look how using the 4 node example above. Suppose "A" goes down and It has a query in its binary log which has to be replicated throught the loop. So the query propagates from "B" to "C" and finally to "D". Now since "A" is down and failover has happened, "B" would be the slave of "D". So the query which originated from "A" would go from "D" to "B". Thats because if "A" would have been there, it would have identified its own query and stopped it. But since "A" is out of the loop, the query will not be stopped and it will propagate to "B" and will always be in the loop. This can result in either the query running in the loop indefinitely or an error on all the slaves and all the slaves going down.

These are the major issues with multi-master replication and automatic failover. Though one can still live with the automatic failover scenario - as it might occur once in a while. But the latency between first and last node during replication has no solution. This latency is due to physical constraints and cannot be avoided.

A work around would be distributing the queries based on tables on all the nodes. So that queries for a table would always be served from a single node. But this then would result in table locks on that node and again we would not be able to reap the benefits of multi-master replication.

Monday, May 07, 2007

Indian railways

It has been almost 27 years since I was born and since i had been traveling in the indian railways. But with increase in time the indian railways has become from bad to worse. It is actually an ordeal to travel in railways. Then why am i writing the post today and not before. Well, i actually realized it yesterday.

At indian railways, the trains are never expected to arrive on time. All trains run late. And if you are standing at one of the platforms, you will constantly hear - "train x which was supposed to come at time y is delayed by delta minutes/hours. The inconvenience caused is deeply regretted". And the announcer would never sound as if he/she was regretting anything. It is said in a matter-of-fact way - as if we are supposed to know that the delay was naturally expected - just like you expect corruption in each and every section of the government of like you expect to get ketchup free with your samosa.

Let me brief you as to what happened last night. Had been to the delhi station to put my mom on the train for baroda. First of all - it had been a long time since i had been to the railways. So i had to enquire about where the parking was. And very few people seemed to know where it was. And since there is no visible board giving directions - i had to run here and there to get the location of parking space.

Well next - get information about the platform where the train would arrive. And the only way to get this information is to stand in queue - after about 100 people and wait your turn. The uncleji at the enquiry window - does not give a shit about who is on the other side of the window. I had time to enquire - what about people who arrived late and need the info - can they wait in queue for 30 minutes to enquire about which platform they should go. It seems the situation in delhi was worse. Its better in baroda - where there is less crowd and the information is properly displayed on the general information board.

Now - since i know which platform i have to go, i will have to spend rs 3/- to get a platform ticket - so that any Ticket Checker cannot catch me and ask for a 50/- or 100/- ghoos for travelling without a ticket. They usually do. There has been so many instances with my friends where the TC simply asked for a Rs 50/- note to let them go. And the way to avoid this hassle is to get a platform ticket. And to get a platform ticket is another hassle. On new delhi station, it is really difficult first to locate the window which gives platform ticket. Officially platform ticket should be available on all windows. But it seems due to scarcity - its there only on one window. And the queue in front of the window is huge. Hmm, there are always queues at all windows on indian railway stations. It seems that the indian railways is very low on manpower. Well actually thats how Laalu prasad yadav - our respected railway minister brought the operating cost of Indian railways to 78% - the best in the world.

And after getting the platform ticket, we proceeded to the platform. Wow, what a place. It seems that people like to live on the platform. Everywhere - there are people sitting, chatting, boozing, gambling. You have to look for space to put your foot to move forward. I think, this is what they mean by population explosion.

What happens in baroda station is that there is a display on the platform which says where each coach will come for the train. So it becomes easy to move and get placed right in front of the coach. But in Delhi - the capital of india, the numbers on the boards are misplaced so that you have to drag your luggage and run in between the crowded platform to your bogie.

And then the trains are always crowded. It is difficult to get a ticket in a train in indian railways. There is always a waiting list and extra people on the train for whom it is important to reach their destination, but were unable to get a reservation. There are people sleeping on the floor, outside the loo and sitting on your berths.

60 years of independence, and indian railways - it seems - stays where it was.

Friday, May 04, 2007

ctrl C & ctrl V

Your Colleague: Hey!! Kya yahan baitha mail forward karta rahta hai yaar!! Naye packages dekh.... Naye language seekh, Night out Maar....Fundoo programming kar like me....! Do something cool man!!

You: Achha! To usse Kya hoga...

You 're Colleague: Impression!! ! Appraisal!!! Har appraisal main tu No 1! Hike in salary!! Extra Stocks

You: Phir kya hoga...

Your Colleague: Project Leader ban jaayega..Phir Project Manager!!! Phir Business Manager! One day U will be a Director of the Company man !!

You: Acchha to phir kya hoga...

Your Colleague: Abe phir tu aish karega! Koi kaam nahin karna padega! Araam se office aayega aur MAIL check karega.

You: To ab main kya kar raha hoon????


"Dikhawe pe na jao, apni akal lagao.

Programming hai waste, trust only copy-paste "


Powered by ctrl C


Driven by ctrl V

Saturday, April 28, 2007

philosophy

I would say that the root cause of all misunderstandings is expectations. It is human nature to expect more than what he is getting now. Either you are not satisfied with your job profile or your salary or your organizations work culture. You expect them to get better. And once they get better, you expect them to get even better. There is no end to the amount of expectations.

It is same in relationships, you expect that your close ones grow up in life. That your wife takes care of you. And your wife would expect you to take care of her, of her emotions, of her will. You expect that your wife or your hubby has great personality in front of others. That others say that you two are a nice couple. You expect that you both listen to each other and respect each other. Etc.. Etc..

The expectations never end, and so neither does the issues or misunderstandings that come out of it.

And the best and simplest way to get out of this is "not to expect". Wow, you would say - how is that possible. How can you not expect. How can you let go of your emotions. How can you say that "let things be as it is".

It is the biggest thing to ask. But that is what the monks do. They just let go of their emotions. Lead a life where they dont care for anything. They meditate. Live with the bare minimum. And they are happy most of the times.

I dont know whether they believe in GOD or not. I dont. How can you believe something that you have not seen, not felt, not heard. Do you believe your friend if he says something odd. I wont. I will have to look at it to believe it. Oh, what would you say if i tell you that i met GOD yesterday. Would you believe me? No way... I think it is the same thing. People say that there is GOD, but where??

And then there is love. I think, i have written about it before. A feeling that cannot be defined. How can you expect something from someone you love. That is not love. That is possessiveness. You dont have expectations in love. All you intend is to give to the one you love. You take care of him/her. You forget and forgive all his/her mistakes. You want that person to be healthy. In all if you love someone, you would expect that person to be happy. And would do anything to make him/her happy. That i think is love.

So expectations dont come in between GOD(whom i dont believe in) and love. But in this materialistic world - you have to eat, you have to drink, and you have to live. And you have some materialistic/physical and emotional needs. And this is where expectations come in picture and they ruin everything.

Whenever you expect something from someone - you open a path to misunderstandings and issues.

Eventually everyone who is born has to die. Then why not make this miserable life happy for others and for self.

LiveJournal - system architecture

Lets discuss the system architecture of Live Journal.

Live Journal or LJ for short kicked off as a hobby project in April 1999 and was built on open source completely. It reached 2.8 Million accounts in April 2004 and 6.8 Million accounts in April 2005. Currently It has more than 10 Million accounts. Caters to several thousands of hits per second and lots of MySQL queries.

Here is a complex diagram which roughly outlines the architecture of LJ.

The technologies which are visible over here are -

Caching - Memcached
Mysql Clusters - HA & LB
Httpd load balancing - using perlbal
MogileFS - Distributed File System

Lets start off with mysql...

A single server with mysql wont be able to handle the large no of reads and writes. With increasing no of reads and writes, the server slows down. Next stage would be to have 2 servers in a master-slave architecture in which the master handles all inserts and the slaves are read-only. But then the queries have to be spread over in such a manner that replication lag between master and slave (though very small) is handled. As the number of database and web servers are increased - chaos increases. Site is fast for a while and then again slow - and there is need for more servers with higher configurations. Also, as the number of slaves increases, the number of writes to the slave also increases. So eventually you come to a situation where the number of writes is very large as compared to the number of reads. Resulting in large I/O and low CPU utilization.

The best way to handle such situation is to divide the database. How LJ did this was by creating user clusters. So each user was assigned a cluster number. And each cluster had multiple machines in a master-slave fashion. The first query would then find the cluster number for that user from the global database and subsequent queries for that user could then be redirected to the user cluster. Ofcourse few issues like uniqueness of userid, and moving user around clusters had to be tackled. Caching of mysql connections and using mysql query cache to cache query results added to the better performance of the site.

Again the problem was the single point of failure with the master databases. If any of the master database dies, the site would go down. To avoid this situation master-master cluster was created. In case of any problem - the other master would come into play and handle all active connections.

Which database engine to use - InnoDB or MyISAM. InnoDB allows concurrent reads and writes and so is comparatively fast. Whereas MyISAM has table level locks and so is not as fast as InnoDB.

And then there is MySQL cluster which is an in-memory engine. It requires about 2-4x of RAM for the dataset. So it is good for handling small data sets only.

An even better way of storing database is by using shared storage - SAN, SCSI, DRDB. You turn a pair of InnoDB machines to a cluster - looks like a single box from outside with floating IP address. Heartbeat to move IP, mount/unmount filesystem, start/stop mysql. DRDB can be used to sync one machine's block device with another. This requires dedicated gigabit cable between the two machines to handle the high amount of data transfer.

Cache

Memcache is used to cache records which has already been computed for frequent access. Memcache is an open source distributed caching system - instances of which can be run on any machine where-ever free memory is available. It also provides simple APIs for different languages like java, php, perl, python and ruby. And it is extremely fast.

LJ created 28 instances of memcache on 12 machines (not dedicated) and was able to cache 30 GB of data. This cache was getting a hit rate of 90-93%. Which reduced the number of queries to the database to a great extent. They started caching stuff which was very frequently accessed and aim at caching almost everything possible. With cache - there is an extra overhead of updating the cache.

http load balancing

After trying a large number of reverse proxies, LJ people were unable to find anything which satisfied their needs. So they built up their own reverse proxy - perlbal - a small, fast, manageable, HTTP web server which can do internal redirects.
It is single threaded, asynchronous and event based. Handles dead nodes. And works in multiple modes - static web server, reverse proxy and plug-ins.

Allows persistent connections and has no complex load balancing logic - uses whatever is free. Connects fast and has multiple queues - for free and paid users.


MogileFS - Distributed File System

Files belong to classes. It tracks what disks are files on. Keeps replicas on devices on different hosts. It has libraries available for most of the languages - php, perl, java, python.

clients, trackers, mysql database cluster and storage nodes - all were brought under MogileFS. It handles automatic file replication, deletion etc.

Have put in only major points and finer details can be found in the link below.


source:
http://danga.com/words/2005_oscon/oscon-2005.pdf

Tuesday, April 24, 2007

realtime fulltext index

For past some days, i have been struggling with getting something which can allow me to create, insert, update and search fulltext indexes - big fulltext indexes. Not very large but somewhere around 3-4 GB of data.

You would suggest mysql, but with mysql the inserts on table with fulltext indexes is very slow. For every insert, data is put in the index which leads to slow inserts.

What else? Well, i had come across some other fulltext engines like sphinx and senna. Compiled mysql with senna but there was no benefit. The index size was almost double and the searches were also slow - as slow as mysql fulltext index.

What about sphinx. Well, had integrated sphinx with mysql. But the sphinx engine has lots of limitations - like the first 2 columns must be integer and the 3rd column should be a text. Rest all columns should be integers. So if i use sphinx, i would need to write a stored procedure and trigger it to port data from my current table to a parallel sphinx table whenever a new row is inserted. And what would i do if i have to run a query - i would be joining both the tables. Would that be fast. Dont know. Let me figure out how to get this thing running...

You would say - what about lucene - my favorite search engine. Well dear, lucene is in java and there is no way i could integrate it with mysql if i have to do it in a jiffy. I would need to design and build a library to get this thing running. Forget it. In fact dbsight provides a similar type of library. Maybe i will get that thing working and check it out. At a small price, I might get what the organization requires and that too without wasting any resources on building it.

Hope i get some solution to the situation i am in right now. I have solutions for it, but not a solution which requires least amount of resources to get it running. Till i get one, the search will continue...

Friday, April 20, 2007

water bridge



Water Bridge in Germany .... What a feat!
Six years, 500 million euros, 918 meters long.......now this is engineering!
This is a channel-bridge over the River Elbe and joins the former East and West Germany, as part of the unification project. It is located in the city of Magdeburg, near Berlin. The photo was taken on the day of inauguration.