Monday, January 7, 2013

About HDInsight




Source: http://www.sqlmag.com/article/sql-server-2012/microsoft-doug-leland-hekaton-hdinsight-sql-server-2012-144898

Otey: Right; this is the new Windows version of Hadoop that you implemented either on-premise[s] or in Azure as a service?

Leland (Microsoft): Yes; so HDInsight for Windows is the on-premise[s] implementation that was announced at Strata. So now our customers broadly have downloadable access to preview bits, for Hadoop on a Windows implementation -- and a preview for the service.
The benefits of these different approaches are kind of like the following. First, we are the only vendor out there who is providing both on-premise[s] offerings and cloud offerings. So the customer has incredible flexibility in their ability both to choose whatever solutions that they need. If I want to run it in my data center, then I'll run it in my data center. If I want to run it in Microsoft's data center, I could now run it in [a] Microsoft data center. I can move it back and forth with the same implementation, 100 percent code compatible. And I can build hybrid scenarios where I may have a piece of the implementation running on data that lives in my data center; or I've got data that's been born in the cloud and I don't really want to bring it back on the prem, so I'll move it into the service. So the flexibility in terms of what we're offering is incredibly valuable to customers in terms of choices that we offer. The flexibility in the scenarios of my data center, the cloud, or some combination of both is also incredibly helpful. And then you kind of bring that back and say, "Now I can use the power of Excel 2013 to get to that data . . . incredibly valuable propositions.
The value of the service, of the HDInsight service, is simpler to the value proposition than you get for products. As a company, I'm not investing in the capital of the infrastructure to build out the software, and I don't have to buy a bunch of machines to build up a cluster. I don't need to hire a Hadoop expert to know how to deploy and build out a Hadoop cluster because I can go to the service, and with basically three clicks and ten minutes, I can deploy a cluster of any size. I pay for that by capacity, but again I don't need that implementation expertise. What I do need is the individual who knows how to build my jobs.


Source: http://www.microsoft.com/web/gallery/install.aspx?appid=HDINSIGHT-PREVIEW

HDInsight is Microsoft's Hadoop-based distribution, built on the Hortonworks Data Platform for Windows (HDP). It combines the simplicity of Windows with the power and reliability of the HDP to deliver big data insights from Apache Hadoop. This is a preview release targeting developer scenarios and as such only support single-node deployments.


-----
Reads:

  1. http://www.microsoft.com/sqlserver/en/us/solutions-technologies/business-intelligence/big-data.aspx
  2. http://hortonworks.com/partners/microsoft/



Friday, January 4, 2013

India - Mera Bharat Mahaan




Pratibha Patil spent Rs 18 crore on her last trip as President, RTI query reveals

Source: http://timesofindia.indiatimes.com/india/Pratibha-Patil-spent-Rs-18-crore-on-her-last-trip-as-President-RTI-query-reveals/articleshow/18007898.cms


NEW DELHI: Notwithstanding a huge controversy over expenditure on her foreign travels, the then President Pratibha Patil ran up a bill of Rs 18.08 crore on her last trip abroad shortly before demitting office, according to official information accessed through RTI.

The chartering of the Air India Boeing 747-400 jumbo for her two-nation trip to South Africa and Seychelles from April 29 to May 8 last year alone cost Rs 16.38 crore, the airline said in an RTI reply.

In addition, an expenditure of Rs 1.46 crore was incurred in Pretoria -- the South African capital. Of this, Rs 71.82 lakh was spent on local stay, 52.33 lakh on transportation and 22.12 lakh on miscellaneous expenditure.

In Durban, an expenditure of Rs 23.55 lakh was incurred. Of this hotel stay alone cost nearly 18 lakh and transportation was Rs 5.27 lakh.

The details about the lodging and other expenditures in South Africa were provided by the Indian Missions in Pretoria and Durban under RTI.

A huge controversy broke out last year when it was revealed that Patil had incurred an expenditure of Rs 205 crore on her 12 foreign trips covering 22 countries across four continents during her five-year term which ended on July 25 last year.

These figures were provided before she undertook the last visit to South Africa and Seychelles.

The Rashtrapati Bhavan had then defended these visits terming them "necessary" to deepen bilateral cooperation. The visits were undertaken after careful appraisal and recommendation by the Prime Minister's office and the ministry of external affairs.



Asaram Bapu rants again, calls media 'barking dogs'

Source: http://www.hindustantimes.com/India-news/NewDelhi/Asaram-Bapu-rants-again-calls-media-barking-dogs/Article1-986670.aspx


A day after blaming the Delhi gangrape victim for the henious crime, Asaram Bapu justified his comments calling those opposing his views as 'barking dogs'. "The media has created a controversy but what wrong did I say? A man complained to me that his wife fights with him; I told him you can't clap with one hand," he said.

"One dog barked and more dogs (the media) joined in. Dogs will bark but they can't harm an elephant's dignity. I didn't intend to harm anybody; I wish the well being of all.."

The family of 23-year-old Delhi gangrape victim has taken a strong exception to ... (sigh, there is more to read)

Delhi rape victim as guilty as her rapists: Asaram Bapu




Spiritual Guru Asaram Bapu has landed himself in a controversy over his remark that the December 16 Delhi [ Images ] gang-rape victim is as guilty as those responsible for the barbaric sexual assault on her.

"Only 5-6 people are not the culprits. The victim daughter is as guilty as her rapists. She should have called the culprits' brothers and begged before them to stop. This could have saved her dignity and life. Can one hand clap? I don't think so," media reports quoted Asaram Bapu, as saying.

According to media reports, the self-proclaimed godman further said that he is against harsh punishments for the accused, as the law could be mis-utilised.

"We have often seen such laws are made to be mis-utilised. Dowry harassment law is the biggest example," he said.

The spiritual guru's remark comes at a time when the entire nation is mourning the death of the 23-year-old girl who died in a Singapore hospital 13 days later after the heinous crime.



Sonia Gandhi travelled in IAF aircraft 49 times in last 7 years

Source: http://news.rediff.com/commentary/2013/jan/04/liveupdates.htm, http://articles.timesofindia.indiatimes.com/2013-01-04/india/36147762_1_iaf-aircraft-air-travel-iaf-rules

15:50 Sonia travelled in IAF aircraft 49 times in last 7 years: UPA Chairperson Sonia Gandhi travelled in Indian Air Force aircraft and helicopters 49 times in the last seven years which included 23 trips with Prime Minister Manmohan Singh.

Congress General Secretary Rahul Gandhi also travelled in IAF aircraft and helicopters eight times in the last three years, according to a reply received under the Right to Information Act.

Both Sonia Gandhi and Rahul Gandhi are not eligible to travel by IAF aircraft and helicopters so they have to travel in the company of eligible persons which includes the Prime Minister, the Deputy Prime Minister, the Home Minister and the Defence Minister for official purposes.

For unofficial purposes, only the Prime Minister is eligible. Other Cabinet ministers can also travel by these aircraft after taking permission of the Prime Minister.



----------
Reads:
  1. http://texolo.blogspot.in/2012/10/open-letter-to-robert-vadra.html


Wednesday, January 2, 2013

Extending laptop battery life


Taken from: http://batterycare.net/en/guide.html

Memory Effect

First of all it's necessary to unfold a myth that persists in many peoples head.
The battery memory effect.
In lithium-based batteries this is in fact a myth, it only applies to older Nickle-based batteries. So fully discharging and charging the battery is completely useless and even harmful as we will see below.
The modern lithium battery can be charged regardless of its current percentage, given that it has absolutely no negative effect in its performance.

Should I remove the battery when A/C is plugged in?

Many laptop users have this question and we will answer it right now:
The answer is: YES and NO, it depends on the situation.
Having a battery fully charged and the laptop plugged in is not harmful, because as soon as the charge level reaches 100% the battery stops receiving charging energy and this energy is bypassed directly to the power supply system of the laptop.
However there's a disadvantage in keeping the battery in its socket when the laptop is plugged in, but only if it's currently suffering from excessive heating caused by the laptop hardware.
So:
- In a normal usage, if the laptop doesn't get too hot (CPU and Hard Disk around 40ºC to 50ºC) the battery should remain in the laptop socket;
- In an intensive usage which leads to a large amount of heat produced (i.e. Games, temperatures above 60ºC) the battery should be removed from the socket in order to prevent unwanted heating.
The heat, among the fact that it has 100% of charge, is the great enemy of the lithium battery and not the plug, as many might think so. 

Battery discharges

Full battery discharges (until laptop power shutdown, 0%) should be avoided, because this stresses the battery a lot and can even damage it. It's recommended to perform partial discharges to capacity levels of 20~30% and frequent charges, instead of performing a full discharging followed by a full charging.
Laptop batteries contain a capacity gauge that allows us to know the exact amount of energy stored. However, due to the charging/discharging cycles, this sensor tends to be inaccurate overtime.
Some laptops include in their BIOS, tools to recalibrate this battery gauge, which is nothing more than a full discharge followed by a full charge.
So to calibrate the gauge, it should be performed, in every 30 discharge cycles, a full discharge non-stop , followed by a also, non-stop, full charge.
An inaccurate gauge can lead to the fact that the the battery capacity values are are wrong. The battery may report that it still has 10% of capacity when in fact it has a much lower value, and this causes the computer to shutdown unexpectedly.
gráfico 3Discharge (or charge) cycles consist of using all that battery charge (100%) but not necessarily all at once.
For example, you can use the laptop for some minutes in a day, using half its capacity e then fully charge it. If you did the same thing in the next day, it would be counted a discharge cycle and not two, so it may take several days until a full discharge cycle is completed.

How to perform a calibration (full discharge)?

The most adequate method to do a full discharge (100% to a minimum of 5%) consists of the following procedure:
  1. Fully charge the battery to its maximum capacity (100%);
  2. Let the battery "rest" fully charged for 2 hours or more in order to cool down from the charging process. You may use the computer normally within this period;
  3. Unplug the power cord and set the computer to hibernate automatically at the minimum percentage possible as described by the image sequence below (click images to enlarge);
            
  4. Leave the computer discharging, non-stop, until it hibernates itself. You may use the computer normally within this period;
  5. When the computer shuts down completely, let it stay in the hibernation state for 5 hours or even more;
  6. Plug the computer to the A/C power to perform a full charge non-stop until its maximum capacity (100%). You may use the computer normally within this period.
After the calibration process, the reported wear level is usually higher than before. This is natural, since it now reports the true current capacity that the battery has to hold charge. Lithium Ion batteries have a limit amount of discharge cycles (generally 200 to 300 cycles) and they will retain less capacity over time.
Many people tend to think "If calibrating gives higher wear level, then it's a bad thing". This is wrong, because like said, the calibration is meant to have your battery report the true capacity it can hold, and it's meant to avoid surprises like, for example, being in the middle of a presentation and suddenly the computer shuts down at 30% of charge.

Prolonged storage

To store a battery for long periods of time, its charge capacity should be around 40% and it should be stored in a place as fresh and dry as possible. A fridge can be used (0ºC  - 10ºC), but only if the battery stays isolated from any humidity.
One must say again that the battery's worst enemy is the heat, so leaving the laptop in the car in a hot summer day is half way to kill the battery.

Purchasing a replacement battery

If you intend to purchase another battery, it's recommended that you do it only when the current battery is very degraded. If it's not the case, the non usage of a battery leads to its degradation.
If a spare battery is purchased and won't be used for a long time, the above storage method should be used.
Besides that, when purchasing a battery you must pay attention to the manufacturing date.

Advantages in using BatteryCare

BatteryCare allows you to have the control over the discharge cycles number, and when this reaches 30 (or other configured value), it notifies you that it's time to perform a full discharge in order to keep the battery gauge calibrated.
Like this, it's guaranteed to always have the correct capacity values reported by the battery.
Besides, when using the battery, there's the possibility to suspend some Operating System features that help degrading the autonomy (only in Windows Vista or higher):
 - Windows Aero, the theme that allows for visual effects like window transparency, requires graphics card acceleration, which obviously will help decreasing the battery lifetime;
- SuperFetch, ReadyBoost and SearchIndexer are three Windows Vista (and higher) services that, even in battery mode, are using the hard disk a lot and increase total power consumption, thus decreasing battery lifetime. Suspending these services has absolutely no negative impact on the performance or security of the system.
These features are resumed once the laptop is plugged in to A/C power.

External links:


Thursday, December 27, 2012

Big data store - HBase


Taken from: http://blogs.igalia.com/dpino/2012/10/31/introduction-to-hbase-and-nosql-systems/



Introduction to HBase and NoSQL systems

October 31st, 2012Go to comment

s

HBase is an open source, non-relational, distributed database modelled after Google’sBigTable (Source: HBase, Wikipedia). BigTable is a data store that relies on GFS (Google Filesystem). Since Hadoop is an open source implementation of GFS and MapReduce, it perfectly made sense to build HBase on top of Hadoop.
Usually HBase is categorized as a NoSQL database. NoSQL is a term often used to refer tonon-relational databases. For instance, Graph databases, Object-Oriented databases, Key-Value data stores or Columnar databases. All of them are NoSQL databases. In the recent years there have been an emerging interest in this type of systems as the relational model has proved to be no effective to solve certain problems, especially those related to storing and handling large amounts of data.
In the year 2000, Berkeley researcher Eric Brewer published a now foundational paper known as the CAP Theorem. This theorem states that it is impossible for a distributed computer system to simultaneously provide all three of the following guarantees:
  • Consistency. All nodes see the same data at the same time.
  • Availability. A guarantee that every request receives a response about whether it was successful or failed.
  • Partition tolerance. The system continues to operate despite arbitrary message loss or failure of part of the system.
According to the theorem, a distributed system can satisfy any two of these guarantees at the same time, but not all three (Source: CAP theorem, Wikipedia).
Usually NoSQL systems are depicted within a triangle representing the CAP Theorem, being each of the angles one the aforementioned guarantees: consistency, availability andpartition tolerance. Each system is located in one of the sides of the triangle, depending on the pair of features favoured.

Visual Guide to NoSQL Systems (Source: http://blog.nahurst.com/visual-guide-to-nosql-systems)
As it is shown in the figure above, HBase is a columnar database that guarantees consistencyof data and partition tolerance. On the other hand , systems like Cassandra or Tokyo Cabinetfavour availability and partition tolerance.
Why NoSQL sytems have become relevant in the recent years? For the last 20 years the storage capacity of hard drives have multiply by several orders of magnitude, however seek and transfer times have not evolved at the same pace. Websites like Twitter receive more data everyday that it can write to a single hard drive, so data has to be written in clusters. Twitter users generate more than 12 TB per day, about 4 PB per year. Other sites, such as Google, Facebook or Netflix, handle similar figures, which means facing the same type of problems. But, big data storage and analysis is not something that only affects websites, for instance, the Large Scale Hadron Collider of Geneva produces about 12 PB per year.
HBase Data Model
HBase is a Columnar data store, also called Tabular data store. The main difference of acolumn-oriented database compared to a row-oriented database (RBMS) is about how data is stored in disk. Check how the following table would be serialized using a row-oriented and a column-oriented approach (Source: Columnar Database, Wikipedia).
EmpIdLastnameFirstnameSalary
1SmithJoe40000
2JonesMary50000
3JohnsonCathy44000
Row-oriented
1,Smith,Joe,40000;
2,Jones,Mary,50000;
3,Johnson,Cathy,44000;
Column-oriented
1,2,3;
Smith,Jones,Johnson;
Joe,Mary,Cathy;
40000,50000,44000;
Physical organization has an impact on features such as partitioning, indexing, caching, views, OLAP cubes, etc. For instance, since common data is stored together, column-oriented excel at operations about aggregating data.
In HBase data is stored in tables, same as in the relational model. Each table consists of rows, each identified by a RowKey. Each row has a fixed number of column families. Each column family can contain a sparse number of columns. Columns also support versioning, that means, that different versions of the same column can exist at the same time. Versioning is usually implemented using a timestamp.
So, to fetch a value from a table we will need to specify three values: <RowKey, ColumnFamily, Timestamp>.
Perhaps the most obscure concept of this model are Column Families. Column Families consist of two parts: a prefix and a qualifier. The prefix is always fixed and it has to be specified when the table is created. However, the qualifier is dynamic and new qualifiers can be added to prefixes at run time. This allows to created an infinite collection of columns inside of column families. Take a look at the table below representing information about Students.
RowKeyTimestampColumnFamily
Student1t1courses:history=”H0112″
Student1t2courses:math=”M0212″
Student2t3courses:history=”H0112″
Student2t4courses:geography=”G0112″
Student2t5courses:geography=”G0212″
It is possible to add new courses just by storing new column families with the prefix courses. To get the code of the history subject Student1 is enrolled in , we need to provide three values: Student1; courses:history; timestamp. It is possible to retrieve all values in case a column family, with a different timestamp, is repeated. In the table above, we can see howStudent2 is enrolled in the Geography course (timestamp=t4, courses:geography=”G0112″, however he enrolled again because the code of the subject changed (timestamp=t5; courses:geography=”G0212″).
Column Families work as a sort of light schema for the tables. Column families have to be specify when a table is created, and it is actually very hard to modify this schema. However, as the qualifier part of a column family is dynamic, it is very easy to add new columns to existingcolumn families.
For further explanation about HBase Data Model I recommend the following articles: Hadoop Wiki -  HBase DataModel, Understanding HBase and BigTable.
Features of HBase
HBase is built on top of Hadoop, that means, it relies on HDFS and it integrates very well with the MapReduce framework. Relying on HDFS provides a series of benefits:
  • A distributed data storage running on top of commodity hardware
  • Redundancy of data
  • Fault-tolerant
In addition, HBase provides other series of benefits (among other features):
  • Random reads and writes (this is not possible with plain Hadoop).
  • Autosharding. Sharding, horizontal data distribution, is done automatically.
  • Automatic failover based on Apache Zookeeper.
  • Linear scaling of capacity. Just add new nodes as you need them.
Installing HBase
HBase depends on Hadoop, so it is necessary to install Hadoop before installing HBase. Currently there are several Hadoop branches being developed at the same time, so it is strongly recommended to install a HBase version that is compatible with a specific version of Hadoop (v1.0, v.0.22, v.0.23, etc). HBase v.0.90.6 is compatible with Hadoop v.1.x. To install HBase follow these steps:
Now, uncompress it:
1
$ sudo tar xfvz hbase-0.90.6.tar.gz
Run HBase:
1
bin/start-hbase.sh
By default, HBase listens on port 60010. When a HBase server is running it is possible to check its state by connecting to http://locahost:60010.
Interacting with the HBase shell
First, start a HBase shell:
1
$ hbase shell
Once you are into a HBase session, create a new table:
1
> create 'students', 'courses'
Now insert some data into the table ‘students’.
1
2
3
4
5
put 'students', 'Student1', 'courses:history', 'H0112'
put 'students', 'Student1', 'courses:math', 'M0212'
put 'students', 'Student2', 'courses:history', 'H0112'
put 'students', 'Student2', 'courses:geography', 'G0112'
put 'students', 'Student2', 'courses:geography', 'G0212'
And now try the command ‘scan’ to show the contents of a table.
1
2
3
4
5
6
7
> scan 'students'
ROW                                    COLUMN+CELL
Student1                              column=courses:history, timestamp=1351332046854, value=H0112
Student1                              column=courses:math, timestamp=1351332046914, value=M0212
Student2                              column=courses:geography, timestamp=1351332047022, value=G0212
Student2                              column=courses:history, timestamp=1351332046950, value=H0112
2 row(s) in 0.0390 seconds
The command ‘scan’ returns two rows (Student1 and Student2). Notice that the column family‘courses:geography’ was inserted twice with different values, but only one value is shown. One of the features of HBase is versioning of data, this means that different versions of the same data can exist in the same table. Generally, this is implemented via a timestamp. So, why those two values don’t show up? To do so, it is necessary to tell scan how many versions of a row we would like to retrieve.
1
2
3
4
5
6
7
8
scan 'students', {VERSIONS => 3}
ROW                                    COLUMN+CELL
Student1                              column=courses:history, timestamp=1351332046854, value=H0112
Student1                              column=courses:math, timestamp=1351332046914, value=M0212
Student2                              column=courses:geography, timestamp=1351332047022, value=G0212
Student2                              column=courses:geography, timestamp=1351332046990, value=G0112
Student2                              column=courses:history, timestamp=1351332046950, value=H0112
2 row(s) in 0.0170 seconds
By default, when a table is created, the number of versions per Column Family is 3. However, it is possible to specify more:
1
> create 'students', {NAME => 'courses', VERSIONS => 10}
To fetch one Student:
1
2
3
4
5
> get 'students', 'Student2'
COLUMN                                 CELL
courses:geography                     timestamp=1351332047022, value=G0212
courses:history                       timestamp=1351332046950, value=H0112
2 row(s) in 0.3100 seconds
The result displays all the current subjects Student2 is enrolled in.
If now we want to disenroll Student2 from subject ‘history’:
1
> delete 'students', 'Student2', 'courses:history'

1
2
3
4
<pre>get 'students', 'Student2'
COLUMN                                 CELL
courses:geography                     timestamp=1351332047022, value=G0212
1 row(s) in 0.0140 seconds
Lastly, we drop table ‘students‘. Dropping a table is two-step operation:
1
2
> disable 'students'
0 row(s) in 2.0420 seconds
And now we can actually drop the table:
1
2
> drop 'students'
0 row(s) in 1.0790 seconds
Summary
This post was a brief introduction to HBase. HBase is a Columnar Database, usually categorized as a NoSQL database. HBase is built on top of Hadoop and shares many concepts with Google’s BigData, mainly its data model. In HBase data is stored in tables, being each table composed of rows and column families. A column family is a pair prefix:qualifier, whereprefix is a fixed part and qualifier is variable. HBase also supports versioning by storing timestamp information for every column family. So, to retrieve a single value from a HBase table, the user has to specify three keys: <RowKey, ColumnFamily, Timestamp>. This way of storing/retrieving data makes that HBase is sometimes referred as a large sparse hash map.
As HBase relies on Hadoop it comes with the common features Hadoop provides: a distributed filesystem, data node fault tolerancy and good integration with the MapReduce framework. In addition, HBase provides other very interesting features such as autosharding, automatic failover, linear scaling of capacity and random reads and writes. The last one is a very interesting feature as plain Hadoop only allows to process the whole dataset in a batch.
And this is all for now. On a following post I will explain how to connect HBase with Hadoop, so it is possible to run MapReduce jobs over data stored in a HBase data store.