Wednesday, July 29, 2020

Linux Sticky Bit Concept Explained with Examples

Linux permissions are a concept that every user becomes intimately familiar with early on in their development. We need to execute scripts, modify files, and run processes in order to administer systems effectively, but what happens when we see Permission denied? Do you know why we see this message? If you know the cause of the problem, do you know how to implement the solution?

I will give a quick explanation of the various ways to calculate permissions, and then we will focus on the special permissions within Linux. If you want an in-depth look at the chmod command, check out this article from Sudoer Shashank Hegde, Linux permissions: An introduction to chmod.

The TL;DR is that there are two main ways of assigning permissions.

Symbolic method

The symbolic method uses the following syntax:

[tcarrigan@server ~]$ chmod WhoWhatWhich file | directory

Where:

  • Who - represents identities: u,g,o,a (user, group, other, all)
  • What - represents actions: +, -, = (add, remove, set exact)
  • Which - represents access levels: r, w, x (read, write, execute)

An example of this is if I want to add the read and write permissions to a file named test.txt for user and group, I use the following command:

[tcarrigan@server ~]$ chmod ug+rw test.txt

Full disclosure, this is not my preferred method of assigning permissions, and if you would like more information around this method, I recommend your nearest search engine.

Numeric method

The numeric method is, in my experience, the best way to learn and practice permissions. It is based on the following syntax:

[tcarrigan@server ~]$ chmod ### file | directory

Here, from left to right, the character # represents an access level. There are three access levels—user, group, and others. To determine what each digit is, we use the following:

  • Start at 0
  • If the read permission should be set, add 4
  • If the write permission should be set, add 2
  • If the execute permission should be set, add 1

This is calculated on a per access level basis. Let's interpret this permissions example:

-rw-r-x---

The permissions are represented as 650. How did I arrive at those numbers?

  • The user's permissions are: rw- or 4+2=6
  • The group's permissions are: r-x or 4+1=5
  • The others's permissions are: --- or 0

To put this into the command syntax, it looks like this:

[tcarrigan@server ~]$ chmod 650 test.txt

Now that you understand the basics of permission calculation in Linux, let's look at the special permissions included in the OS.

Special permission explained

Special permissions make up a fourth access level in addition to user, group, and other. Special permissions allow for additional privileges over the standard permission sets (as the name suggests). There is a special permission option for each access level discussed previously. Let's take a look at each one individually, beginning with Set UID:

user + s (pecial)

Commonly noted as SUID, the special permission for the user access level has a single function: A file with SUID always executes as the user who owns the file, regardless of the user passing the command. If the file owner doesn't have execute permissions, then use an uppercase S here.

Now, to see this in a practical light, let's look at the /usr/bin/passwd command. This command, by default, has the SUID permission set:

[tcarrigan@server ~]$ ls -l /usr/bin/passwd 
-rwsr-xr-x. 1 root root 33544 Dec 13  2019 /usr/bin/passwd

Note the s where x would usually indicate execute permissions for the user.

group + s (pecial)

Commonly noted as SGID, this special permission has a couple of functions:

  • If set on a file, it allows the file to be executed as the group that owns the file (similar to SUID)
  • If set on a directory, any files created in the directory will have their group ownership set to that of the directory owner
[tcarrigan@server article_submissions]$ ls -l 
total 0
drwxrws---. 2 tcarrigan tcarrigan  69 Apr  7 11:31 my_articles

This permission set is noted by a lowercase s where the x would normally indicate execute privileges for the group. It is also especially useful for directories that are often used in collaborative efforts between members of a group. Any member of the group can access any new file. This applies to the execution of files, as well. SGID is very powerful when utilized properly.

As noted previously for SUID, if the owning group does not have execute permissions, then an uppercase S is used.

other + t (sticky)

The last special permission has been dubbed the "sticky bit." This permission does not affect individual files. However, at the directory level, it restricts file deletion. Only the owner (and root) of a file can remove the file within that directory. A common example of this is the /tmp directory:

[tcarrigan@server article_submissions]$ ls -ld /tmp/
drwxrwxrwt. 15 root root 4096 Sep 22 15:28 /tmp/

The permission set is noted by the lowercase t, where the x would normally indicate the execute privilege.

Setting special permissions

To set special permissions on a file or directory, you can utilize either of the two methods outlined for standard permissions above: Symbolic or numerical.

Let's assume that we want to set SGID on the directory community_content.

To do this using the symbolic method, we do the following:

[tcarrigan@server article_submissions]$ chmod g+s community_content/

Using the numerical method, we need to pass a fourth, preceding digit in our chmod command. The digit used is calculated similarly to the standard permission digits:

  • Start at 0
  • SUID = 4
  • SGID = 2
  • Sticky = 1

The syntax is:

[tcarrigan@server ~]$ chmod X### file | directory

Where X is the special permissions digit.

Here is the command to set SGID on community_content using the numerical method:

[tcarrigan@server article_submissions]$ chmod 2770 community_content/
[tcarrigan@server article_submissions]$ ls -ld community_content/
drwxrws---. 2 tcarrigan tcarrigan 113 Apr  7 11:32 community_content/

Summary

In closing, permissions are fundamentally important to being an effective Linux administrator. There are two defined ways to set permissions using the chmod command: Symbolic and numerical. We examined the syntax and calculations required for both methods. We also considered the special permissions and their role in the system. Now that you understand permissions and the underlying concepts, you can solve the ever-annoying Permission denied error when it tries to impede your work.





Wednesday, November 27, 2019

Remove all from docker

1. Remove all in docker.
# Stop all containers
docker stop `docker ps -qa`

# Remove all containers
docker rm `docker ps -qa`

# Remove all images
docker rmi -f `docker images -qa`

# Remove all volumes
docker volume rm $(docker volume ls -qf dangling="true")

# Remove all networks
docker network rm `docker network ls -q`

# Your installation should now be all fresh and clean.

# The following commands should not output any items:
# docker ps -a
# docker images -a
# docker volume ls

# The following command show only show the default networks:
# docker network ls

source https://gist.github.com/beeman/aca41f3ebd2bf5efbd9d7fef09eac54d

2. Debug in docker
# method 1
docker run -it image_id bash

Sunday, May 19, 2019

5 Most Common Python Interview Questions with Detailed Answers

1. What is pickling and unpickling?

Pickling means serializing the object into binary streams before writing it to the file. It’s basically a way to convert a Python object into a character stream.
In Python, any object can be pickled to be saved on disk. This is very useful when someone wants to save the state of his objects and reuse later without losing any instance-specific data.
In simple words, Pickling and unpickling are used for serializing and deserializing an object structure in Python.
Suppose, you’re restoring a saved game, it means you’re loading pickled, and therefore you’re unpickling it. Both of these terms are widely used in the industry as they allow you to easily send data from one server to another and ultimately store it in a file or database.

2. What is a Tuple? What’s the difference between lists and tuples in Python?

Both lists and tuples are sequence data types that can store a collection of items. But the main difference between these two is — lists are mutable and tuples are immutable. It means you can modify the values in lists but you cannot modify or even copy the values in tuples.
Another difference is lists are used to store homogeneous elements i.e elements belonging to the same type, whereas tuples are used to store heterogeneous elements i.e elements belonging to different types. You can also use tuples as dictionary keys if needed.

3. What are some advantages of using Python over other programming languages?

There are a lot of advantages to using Python over other languages. First of all, Python is open-source, has extensive support libraries and clean object-oriented designs. It has simple syntax and comes with easy readability.
Python has powerful control capabilities. It can call C, C++ or Java directly via Jython. It comes with Enterprise Application Integration which makes it easier to develop Web services by invoking COM or COBRA components. Python also saves programming time and effort by reducing the number of lines of code required to do a particular task.

4. How memory is managed in Python?

Python contains a private heap storing all the Python objects and data structures. This private heap is internally managed by the Python memory manager.
This Python memory manager has different components which are responsible for various dynamic storage management aspects such as sharing, segmentation, caching or preallocation.
Hence, the Python memory manager delegates some of its work to the object-specific allocators but always ensures that they operate within the private heap limits.

5. What is Flask and what are its benefits?

Flask is a Python framework inspired by Sinatra Ruby framework and based on the Werkzeug and Jinja2 framework libraries.
The main purpose of building this framework was to create a strong foundation for both simple and complex web applications. Flask is easy to use and extend. It allows one to use any extension he/she needs.
They are Unicode based, have a built-in development server, support for secure cookies, has integrated support for unit testing. Flask is available under the BSD license.

Thursday, May 9, 2019

sed, awk command line

1. Insert text from begining and special line (from line >= 57) with awk
  • awk 'FNR>=57,$0="allow "$0{print}' /etc/nginx/akamai_https.conf > test2.txt
  • sed -i -e 's/^/text\ to \ insert/g' filename
2. Replace all IP "xxx.xxx.xxx.xxx;" => "xxx.xxx.xxx.0/24;"
  • sed -e 's/\.[0-9]\+;/\.0\/24;/g' /tmp/ips.txt > /tmp/new_ips.txt
    Here is content in file:
    access_log /var/log/test.log;
    I want to insert before the ; character
  • sed -i -e "s/\(access_log.*.log\)\(;\)/\1\ combined\ if\=\$loggable\2/g" filename.txt
3. Replace string in all file in dir
  • ls -1 <dir> | xargs -n1 -i%f sed -i -e 's/old_string/new_string/g' '%f'
4. Sed with variable with forward slash in shell script
  • vari1=/Users/huytn/test
  • sed -i -e "s/old_string/${vari1//\//\\/}/g" /path/to/file







Tuesday, February 12, 2019

10 thứ trẻ nên học để trở thành người hạnh phúc

1. Ngoại ngữ
Một nghiên cứu cho thấy những đứa trẻ học ngôn ngữ thứ hai thích nghi nhanh với sự thay đổi, có trí nhớ tốt và hiểu rõ hơn về ngôn ngữ nói chung. Chưa kể, khi lớn lên, việc giao tiếp được bằng nhiều thứ tiếng sẽ giúp con bạn có nhiều lựa chọn nghề nghiệp hơn. 
Các nhà khoa học từ Viện nghiên cứu Rotman ở Canada cũng đã chứng minh việc nói hai ngôn ngữ giúp não trì hoãn sự khởi phát của bệnh Alzheimer khi về già.
2. Bơi lội
Hoạt động thể chất giúp chúng ta có cuộc sống lành mạnh. Bơi lội còn là kỹ năng sinh tồn cần thiết giúp trẻ tự cứu sống bản thân trong những tình huống nguy cấp. Nhờ bơi lội, tay chân sẽ phát triển khả năng phối hợp, não bộ duy trì sự minh mẫn, theo nghiên cứu được công bố bởi Trung tâm Thông tin Công nghệ sinh học Quốc gia Mỹ.
3. Nhạc cụ
Ảnh: Pixabay
Ảnh: Pixabay
Theo Journal of Neuroscience, việc chơi một nhạc cụ giúp cải thiện kỹ năng của thính giác, trì hoãn sự suy giảm năng lực não bộ khi về già. Điều này xảy ra do khi chơi nhạc cụ, chúng ta kích hoạt một số hệ thống ở não cùng lúc như thính giác, vận động và nhận thức. Nhờ chơi nhạc cụ từ nhỏ, trẻ sẽ có khả năng giao tiếp và thể hiện bản thân tốt hơn khi lớn lên. 
4. Nhảy múa
Một nghiên cứu từ Đại học Karlstad (Thụy Điển) cho thấy nhảy múa giúp trẻ học cách giao tiếp và thể hiện cảm xúc thông qua cơ thể. Âm nhạc còn khuyến khích khả năng sáng tạo, kỹ năng xã hội và vận động. Nhảy múa giúp trẻ đến gần hơn với những nền văn hóa khác, khiến chúng có xu hướng cởi mở hơn và tự tin vào cơ thể của mình. 
5. Tái chế
Thông qua việc tái chế, chúng ta góp phần bảo vệ Trái Đất, để lại một thế giới tốt hơn cho thế hệ sau. Đối với trẻ em, tái chế còn kích thích khả năng sáng tạo. Chúng sẽ biết không cần đến những nguyên liệu đắt tiền để biến ý tưởng thành hiện thực. Ở Tây Ban Nha, nhiều nhà giáo dục khuyến khích đưa tái chế vào giảng dạy trong trường học để nâng cao nhận thức về bảo vệ môi trường. 
6. Dọn dẹp
Ảnh: Depositphotos
Ảnh: Depositphotos
Trật tự và vệ sinh là hai yếu tố không thể thiếu trong cuộc sống của bất kỳ người nào. Bên cạnh những lý do dễ thấy như mang lại không gian sạch đẹp, việc dọn dẹp còn giúp ích cho tinh thần, khiến con người trở nên có tổ chức hơn. Ở Nhật Bản, việc dọn dẹp lớp học và khuôn viên trường là một phần của giáo dục. 
7. Định hướng
Khuyến khích ý thức định hướng sẽ tốt cho não bộ của trẻ. Một nghiên cứu đã chỉ ra rằng có một hệ thống định vị bên trong não bộ tạo ra các mạng lưới tế bào thần kinh, nuôi dưỡng ý thức và giúp não lập kế hoạch, lộ trình, cải thiện khả năng đưa ra quyết định của mỗi người.

8. Nấu ăn
Bằng cách nấu ăn, trẻ sẽ cải thiện mối quan hệ với thực phẩm. Nếu bạn cho con cùng chuẩn bị một món ăn, chúng sẽ thích thú hơn khi ngồi vào bàn ăn và dần ít muốn ăn vặt. Khi làm theo công thức nấu ăn, chúng sẽ học được tầm quan trọng của việc tuân thủ hướng dẫn, khám phá từng nguyên liệu bằng các giác quan cụ thể. Nếu trẻ còn quá nhỏ, bạn hãy giao nhiệm vụ đơn giản, quan sát kỹ để tránh xảy ra tai nạn. 
9. Tiêu tiền
Khi lớn lên, trách nhiệm tài chính ngày càng nặng nề và chúng ta sẽ dễ mắc sai lầm nếu không được chuẩn bị từ sớm. Điều quan trọng là hãy dạy trẻ rằng tiền là công cụ, không phải phần thưởng. Không chỉ học cách tiết kiệm, trẻ cũng cần học cách tiêu tiền khôn ngoan. 
10. Biểu lộ cảm xúc
Ảnh: Pixabay
Ảnh: Pixabay
Một số tình huống không tránh được trong cuộc sống gây cảm xúc tiêu cực. Do đó, ngay từ khi con còn nhỏ, bạn hãy dạy chúng cách xác định, chấp nhận và biểu lộ cảm xúc. Trí tuệ cảm xúc (EQ) sẽ cho phép trẻ đưa ra quyết định và phản ứng phù hợp trong các tình huống phức tạp. 
Ngoài ra, bạn nên dạy con tầm quan trọng của việc nghỉ ngơi và thư giãn. Ở trường học, thầy cô dạy trẻ kiến thức, giao bài tập, yêu cầu thu thập thông tin và cách sống hòa hợp với tập thể. Ở nhà, cha mẹ cần giúp trẻ học cách tìm những khoảnh khắc thích hợp để vui đùa và giải phóng sự căng thẳng. 

Tuesday, December 18, 2018

gitlab some tips


  • Working with non default ssh key pairs path
    • If you use non default file path for your Gitlab key pairs, you must configure your ssh client to find your gitlab private ssh key for connections to gitlab. Open terminate and enter
$eval($ssh-agent -s)
$ssh-add /path/to/another_id_rsa
    •  To retain these settings, you need to save them to configuration file. For openssh client this is configured in the ~/.ssh/config file
Host gitlab.com
    Preferredauthentications publickey 
    IdentityFile ~/.ssh/example_com_rsa 
  • Switching remote URLs from https to git: 
    • open terminal
git remote set-url https://gitlab.com gitlab@gitlab.com:USERNAME/REPOSITORY
    • Verify  remote URL has changed
git remote -v


Thursday, June 28, 2018

About Secure Password Hashing

An often overlooked and misunderstood concept in application development is the one involving secure hashing of passwords. We have evolved from plain text password storage, to hashing a password, to appending salts and now even this is not considered adequate anymore. In this post I will discuss what hashing is, what salts and peppers are and which algorithms are to be used and which are to be avoided.

Hashing

Hashing is a type of algorithm which takes any size of data and turns it into a fixed-length of data. This is often used to ease the retrieval of data as you can shorten large amounts of data to a shorter string (which is easier to compare). For instance let’s say you have a DNA sample of a person, this would consist of a large amount of data (about 2.2 – 3.5 MB), and you would like to find out to who this DNA sample belongs to. You could take all samples and compare 2.2 MB of data to all DNA samples in the database, but comparing 2.2 MB against 2.2 MB of data cant take quite a while, especially when you need to traverse thousands of samples. This is where hashing can come in handy, instead of comparing the data, you calculate the hash of this data (in reality, several hashes will be calculated for the different locations on the chromosomes, but for the sake of the example let’s assume it’s one hash), which will return a fixed length value of, for instance, 128 bits. It will be easier and faster to query a database for 128-bits than for 2.2 MB of data.
The main difference between hashing and encryption is that a hash is not reversible. When we are talking about cryptographic hash functions, we are referring to hash functions which have these properties:
  • It is easy to compute the hash value for any given message.
  • It is infeasible to generate a message that has a given hash.
  • It is infeasible to modify a message without changing the hash.
  • It is infeasible to find two different messages with the same hash.
The hash function should be resistant against these properties:
  • Collisions (two different messages generating the same hash)
  • Pre-image resistance: Given a hash h it should be difficult to find any message m such that h = hash(m).
  • Resistance to second-preimages: given m, it is infeasible to find m’ distinct from m and such that MD-5(m) = MD-5(m’).

Modern Hashing Algorithms

Some hashing algorithms you may encounter are:
  • MD-5
  • SHA-1
  • SHA-2
  • SHA-3

MD-5

MD-5 is a hashing algorithm which is still widely used but cryptographically flawed as it’s prone to collisions. MD-5 is broken in regard to collisions, but not in regard of preimages or second-preimages. The first attacks on MD-5 were published in 1996, this was in fact an attack on the compression of MD-5 rather than MD-5 itself. In 2004 a theoretical attack was produced which allowed for weakening the pre-image resistance property of MD-5. In practice the attack is way too slow to be useful.

SHA

SHA or Secure Hashing Algorithm is a family of cryptographic hash functions published by the National Institute of Standards and Technology (NIST) as a U.S. Federal Information Processing Standard (FIPS). Currently three algorithms are defined:
  • SHA-1: A 160-bit hash function which resembles the earlier MD-5 algorithm. This was designed by the National Security Agency (NSA) to be part of the Digital Signature Algorithm. Cryptographic weaknesses were discovered in SHA-1, and the standard was no longer approved for most cryptographic uses after 2010.
  • SHA-2: A family of two similar hash functions, with different block sizes, known as SHA-256 and SHA-512. They differ in the word size; SHA-256 uses 32-bit words where SHA-512 uses 64-bit words. There are also truncated versions of each standardized, known as SHA-224 and SHA-384. These were also designed by the NSA.
  • SHA-3: SHA-3 is not yet defined. NIST is working on the exact parameters they will use; SHA-3 will be Keccak, or “close enough”, but not necessarily the Keccak which was submitted (it is a configurable function, and they seem to want to tweak the parameters a bit differently than what was first proposed).
Note that while SHA-1 is “cryptographically broken” the properties we seek in a password hashing algorithm are still valid. In the real world finding a password hashing algorithm built on SHA-1 is still secure in the sense, that if it’s implemented there is no reason to assume it should be immediately changed to something newer.

Strong passwords

Apart from choosing a good hashing algorithm you should also force your users to choose a password which is built up of at least eight, random characters. Unfortunately people aren’t designed to remember and generate random sequences of characters. This is why we force our users to make passwords which contain numbers, letters, signs and at least one capital letter. But how does this help in regard to password hashing?
To attack hashed passwords there are different strategies:
  • Dictionary Attacks
  • Bruteforce
  • Rainbow Tables (generate everything upfront in a database and do a look up for each hash)
With a dictionary attack you will try to use word lists, these can consist of mostly used passwords, words, names, years, etc. For each word you will run the hashing algorithm and see if the generated hash is the same as the hash in the database. If this is the case then you know that the word from which you derived the hash is the password.
With a bruteforce attack you will try all possible combinations of characters. When using passwords of at least eight characters long, only using the ASCII characters set, there are 128^8 possibilities of passwords.
To show the importance of the length of a password:
These days, using a single, modern GPU, you can process about 10.323.000.000 passwords per second when bruteforcing plain MD-5. With this speed, when using a password of eight random characters, it will take about eighty days to generate every single possibility. This single GPU only costs about 500 USD (AMD Radeon 6990). People have actually constructed clusters which contain 25 of these cards, optimized it and managed to generate 350 billion passwords per second. This means they can generate all possible passwords of eight random characters long in less than two days.
Now when you add one character to the password, the possibilities will be 128^9. With previous calculation of 350 billion it will now take 305 days. 10 characters -> 106 years. This seems long, but we need to take into account Moore’s law:
Moore’s law is the observation that, over the history of computing hardware, the number of transistors on integrated circuits doubles approximately every two years. The period often quoted as “18 months” is due to Intel executive David House, who predicted that period for a doubling in chip performance (being a combination of the effect of more transistors and their being faster).
Computers have become faster and faster over the years, which is something we need to take into account. From a cryptographical point of view, 106 years is still a short period. We want infinity (something which will take several hundred-thousand to millions of years).

Hashing Passwords

Why do we hash passwords?

We hash passwords because in the event an attacker gets read access to our database, we do not want him to retrieve the passwords plain text. Remember often we store usernames, email addresses and other personal information in our database. Security rule #1 dictates that users need to be protected from themselves. We can make them aware of the risks, we can tell them not to re-use passwords, but we all know that in the end there will still be people who use the same password for their Facebook, Gmail, Linkedin and corporate email. What you do not want is that when the attacker gets his hand on your database, he immediately has access to all the above accounts (usernames/email addresses will be the same).
Hashing passwords is to prevent this from happening, when the attacker gets his hands on your database, you want to make it as painful as possible to retrieve those passwords using a brute-force attack. Hashing passwords will not make your site any more secure, but it will perform damage containment in the event of a breach.

Properties

Now that you have a small overview of hashing algorithms let’s dive into password hashing. Password hashing requires the following properties:
  • Have a unique salt per password (salt may only be used once across the database containing all password hashes) to prevent a bruteforce attack of compromising the data in one run
  • Fast on software (executing the function once must be relatively fast)
  • Slow on hardware (executing the function concurrently should be slow, this to prevent brute forcing on distributed systems)

Salt and Pepper

A salt is a non-secret, unique value in the database which is appended (depending on the used algorithm) to the password before it gets hashed. Note that the only requirement of a salt should be that it is unique in the database, which means that random generation does not need to be cryptographically random. A salt is used to prevent Rainbow Table lookups (an attack where you calculate a table of hashes for passwords).
It should be noted that you should never use any user supplied information as salt. This to prevent that two databases (of another application for example) would return the same hash for a given password+salt. Also when bruteforcing, you can’t generate all possibilities in one run, you will need to run the bruteforce for each password taking the salt into account. This means that instead of being able to break all passwords in 2 days (as mentioned in previous paragraph), you will need to run the bruteforce process for each password taking into account the specific salt for that password. So this means it will take 2 days for every hash (1 day if you take probability into account) rather than 2 days for the complete database. It will also prevent people from using Rainbow Tables, as the table would need to be generated, per password, based on the salt. This is why salts are so important and why they should be unique for a database!
A pepper is a secret value (a key) which is used to turn the hash into a HMAC (The pepper does not necessarily mean HMAC. The pepper is best described as a secret key which turns the hash function into a MAC. There are good and bad ways to do a MAC; one good way is HMAC.). A HMAC is like a hashing function, but cannot be reproduced without knowing the key. This could increase security if an attacker would only have access to the database and not to the place where the key is actually stored. The problem is that often in data compromises an attacker will at least have access to the hard disk and in some cases even the memory. If the attacker can obtain the key, then the pepper does not add any security to the hashing algorithm.

Speed

A password hashing algorithm should be slow to prevent bruteforce attacks, pereferably it should have features which actually decrease the feasibility of a distributed brute force attack on the hashes. This immediately throws out the following hashing algorithms:
  • MD-5
  • SHA-1
  • SHA-2
  • SHA-3
In their normal form all of the above hashing algorithms are actually unsuited for password hashing. They are all incredibly fast. People often think MD-5 is flawed for password hashing because of collision attacks, but this is not true, it’s because it’s an incredibly fast hashing algorithm. The same is true for SHA-3. I’ve seen people use SHA-3 for password hashing. They assume because it’s supposed to become the new standard that it’s suited for password hashing. Keccak (the name of the algorithm) was, by design, meant to be a fast hashing algorithm. This means it’s completely unsuited for password hashing in any form. So please note that not all hashing algorithms are suited for password hashing.
We can also add:
  • MD-5, SHA-1, SHA-256 (and SHA-224) use 32-bit operations, which are very fast on x86 CPU but also on ARM and, crucially, on GPU as they exist today.
  • SHA-512 (and SHA-384) use 64-bit arithmetic operations which make life harder for GPU, but also for 32-bit architectures (ARM in particular).
  • SHA-3 (Keccak) uses only boolean operations, which makes it fast everywhere but also quite faster on FPGA (no carry propagation).
So what to do now? Well you still have a few options open if you want to use the hashes from the Secure Hashing Algorithm suite. But they need to be used in a PBKDF2 implementation (and even then not every one of them is suited), which we will discuss below.

Hashing Passwords: algorithms

There are currently three algorithms which are safe to use:
  • PBKDF2
  • bcrypt
  • scrypt

PBKDF2

PBKDF2 is an algorithm which is used to derive keys. It wasn’t intended for password hashing, but due to it’s property of being slow, it lends itself quite well for this purpose. The resulting derived key (HMAC) can actually be used to securely store passwords. It’s not the ideal function for password hashing, but it’s easy to implement and it’s built upon SHA-1 or SHA-2 hashing algorithms (any HMAC will do, but these are the most common used ones and the most secure ones). Wait, didn’t you say SHA-1 and SHA-2 were bad to use when hashing passwords? Yes indeed, that’s why we use PBKDF2 to make the hashing a lot slower. You still will need to choose your hashing algorithm carefully, PBKDF2+Keccak is a substantially worse choice than PBKDF2+SHA-256, which is already somewhat worse than PBKDF2+SHA-512 if your server is a 64-bit PC.
To derive a key PBKDF2 does the following:
DK = PBKDF2(PRF, Password, Salt, c, dkLen)
Where DK is the derived key, PRF is the preferred HMAC function (this can be a SHA-1/2 HMAC, the password is used as a key for the HMAC and the salt as text), c is the amount of iterations and dkLen is the length of the derived key. A salt should, by definition of the standard, be at least 64-bits of length and the minimum amount of iterations should be 1024. What the algorithm will do is SHA-1-HMAC(password+salt), and repeat the calculation 1024 times on the result. This means the hashing of a password will be 1024 times slower. Still this does not actually offer a lot of protection when bruteforcing on distributed systems or GPU (Graphic Processing Unit).
There’s also a caveat when the password exceeds 64 bytes, the password will be shortened by applying a hash to it by the PBKDF2 algorithm so it does not exceed the block size. For instance when using HMAC-SHA-1 a password longer than 64 bytes will be reduced to SHA-1(password), which is 20 bytes in size. This means passwords longer than 64 bytes do not provide additional security when it comes to breaking the key used to make the HMAC, but may even reduce security as the length of the key will be reduced (note that even when reduced to 20 bytes, currently our great-great-great-great-great-great-great-great-great-great grand children will be long dead before the key is brute forced).

bcrypt

bcrypt is currently the defacto secure standard for password hashing. It’s derived from the Blowfish block cipher which, to generate the hash, uses look up tables which are initiated in memory. This means a certain amount of memory space needs to be used before a hash can be generated. This can be done on CPU, but when using the power of GPU it will become a lot more cumbersome due to memory restrictions. Bcrypt has been around for 14 years, based on a cipher which has been around for over 20 years. It’s been well vetted and tested and hence considered the standard for password hashing.
There is actually one weakness, FPGA processing units. When bcrypt was originally developed it’s main threat was custom ASICs specifically built to attack hash functions. These days those ASICs would be GPUs (password bruteforcing can actually still run on GPU, but not in full parallelism) which are cheap to purchase and are ideal for multithreaded processes such as password bruteforcing.
FPGAs (Field Programmable Gate Arrays) are similar to GPUs but the memory management is very different. On these chips bruteforcing bcrypt can be done more efficiently than on GPUs, but if you have a long enough password it will still be unfeasable.

scrypt

For password hashing, the current fashion is to move the problem away to another level; instead of doing a lot of hash function invocations, concentrate on an operation which is hard for anything else than a PC, e.g. random memory accesses. That’s what scrypt is about. Scrypt is another hashing algorithm which has the same properties as bcrypt, except that when you increase rounds, it exponentially increases calculation time and memory space required to generate the hash. Scrypt was created as response to evolving attacks on bcrypt and is completely unfeasable when using FPGAs or GPUs due to memory constraints. Scrypt requires the storage of a series of intermediate state data “snapshots”, which are used in further derivation operations. These snapshots, stored in memory, grow exponentially compared when rounds increase. So adding a round, will make it exponentially harder to brute force the password. Scrypt is still relatively new compared to bcrypt and has only been around for a couple of years, which makes it less vetted than bcrypt.

Conclusion and Acknowledgments

Passwords should be hashed with either PBKDF2, bcrypt or scrypt, MD-5 and SHA-3 should never be used for password hashing and SHA-1/2(password+salt) are a big no-no as well. Currently the most vetted hashing algorithm providing most security is bcrypt. PBKDF2 isn’t bad either, but if you can use bcrypt you should. Scrypt, while still considered very secure, hasn’t been around for a long time, so it doesn’t get recommended a lot, but it seems it will become the successor of bcrypt, once it has been around a bit longer. Note that while there are some caveats and password bruteforcing strategies for PBKDF2 and bcrypt, they are still considered unfeasable for strong passwords (passwords longer than 8 characters, containing numbers, letters, signs and at least one capital letter).
I would also like to thank Thomas Pornin for supplying so much insights, reviewing my cryptography assumptions and for suggesting amendments. I would also like to thank CodesInChaos for answering my questions regarding PBKDF2 key generation.

https://security.blogoverflow.com/2013/09/about-secure-password-hashing/