Make Your Java Applications Run Faster - Part 3 - Variables and Arrays

In the previous two parts of the series "Make Your Java Applications Run Faster" we saw how the compiler does some of the optimizations for us and some general techniques that can be applied to our code.

In this section, we shall see how usage of arrays and built in variables can be optimized.

Array Initializations: For optimizing your java code, It is important to know how arrays are initialized. Arrays are initialized at run time one element at a time.
For Example:

int[][] positions = {{0,0},{1,2},{2,5}};

The above array initialization is translated literally by the compiler to the equivalent:
int[][] positions = new int[3][];
positions[0]=new int[2];
positions[0][0] = 0;
positions[0][1] = 0;
positions[1]=new int[2];
positions[1][0] = 1;
positions[1][1] = 2;
positions[2]=new int[2];
positions[2][0] = 2;
positions[2][1] = 5;


This sort of initialization creates the following problems:
1) It can bloat your class files by adding more byte code to the class file.
2) When such array is initialized as a local variable, these se of codes will get executed each time the method is invoked.
It is better to declare the array as a static or instance member to eliminate the initialization for each iteration. Even if you need a fresh copy each time the method is called, for non-trivial arrays it is faster to store a single initialized copy and make a copy of it.

Copying Arrays: You can create a copy of an array efficiently by using the System.arraycopy() method. It is a special purpose optimized native method which is faster and efficient than copying each element in a loop.

Care when using byte, short and char: When stored in a variable, built in types like byte, short, char and boolean all are represented as 32 bit values just like int and also use the same bytecode for loading and storing. The differerence in the memory used comes only when these built in types are stored in arrays. When stored in arrays, boolean and byte values are stored as 8-bit values, while short and char values are stored as 16-bit values.
Note that int, float, and object reference values are always stored as 32-bits each, and long and double values are always stored as 64-bit values.
The fastest types of variables are int and reference variables. This is because all operations on byte, short, and char are performed as ints, and assigning the results to a variable of the corresponding type requires an explicit cast.

In the following, (b + c) is an int value and has to be cast to assign to a byte variable:
byte a, b, c;
a = (byte) (b + c);

Casting requires extra bytecode instructions and thus adds additional overhead when using these smaller types namely byte, short and char.
Thus, the benefit of using smaller types comes only when they are stored in and array and are seldom used in arithmetic operations.

Make your Java Application run faster - Part 2 - General Optimization Techniques

In the last part of the series “Compiler Optimizations” we saw some common optimizations done by the java compiler. You need to be aware of these optimizations because the compiler does most of the optimizations for you. In this part of the series, we shall see some optimizations that you can do to your code. These optimizations are general and apply to any programming language.

Strength Reduction: Using cheaper/faster operations in place of expensive ones.

For example, use of the compound assignment operators such as a+=b instead of a=a+b will be faster because they result in fewer byte code instructions.

You can also use shifts instead of multiplication by powers of two, for example, x >> 2 can be used in place of x / 4, and x <<>

Common Sub Expression Elimination: Eliminate redundant calculations in your code by performing those computations only once and storing the result in a temporary variable, then replacing the computations with the temporary variables

For Example: Instead of writing

double p = d * (average / total) * scorep;

double q = d * (average / total) * scoreq;

The common sub expression can be calculated once and used for both calculations:

double depth = d * (average / total);

double p = depth * scorep;

double q = depth * scoreq;

Code Motion: It is similar to sub expression elimination. Here, expensive operations can be moved out of a loop so that it is evaluated only once. We had seen this in compiler level optimizations also, but compiler can perform code motion only for simple expressions- involving only local and final static variables. For example in

for (int i = 0; i <>

x[i] *= computeSinVar(y);

We can move computeSinVar(y) out of the loop because y does not vary inside the loop.

Double sinVar = computeSinVar (y);

for (int i = 0; i <>

x[i] *= sinVar;

Compiler cannot do this optimization automatically because it does not know whether the array x is being modified in the method computeSinVar.

Un Rolling Loops: Reduce the number of iterations of a loop by performing more operations in each iteration of the loop.

The idea is to save time by reducing the number of overhead instructions that the computer has to execute in a loop.

To achieve this, the instructions that are called in each iteration of the loop are replicated to form a single sequence.

There will be significant performance gain only if the operations done in the loop consists of only simple statements such that the loop control consumes a significant part of each iteration.

For Example:- The loop in the previous example can be unrolled as below.

Double sinVar = computeSinVar (y);

for (int i = 0; i <>

{

x[i] *= sinVar;

x[i+1] *= sinVar;

}


In practice, unrolling loops such as this -- in which the value of the loop index is used within the loop and must be separately incremented -- does not yield an appreciable speed increase in interpreted Java because the bytecodes lack instructions to efficiently combine the "+1" into the array index. Moreover, in the above example the loop will work fine only if the array x has even number of elements.

Note that some programming language compilers like c unroll loops as a part of the compiler optimization itself.

Make your Java Application run faster - Part 1 - Compiler Optimizations

It is a general concern that java applications run slower that C and C++ applications. Which this concern is partly true; there are some things that you can do to make your applications faster.

In this series of I will be presenting several articles on optimizing your java applications and suggest several techniques to make them run faster.

In this Part we shall see the various Compiler Level Optimizations.

Before going on to the changing your code, it is better to know the optimization that the compiler is already doing. Java compiler is intelligent and does many optimizations that even if you write sloppy code the compiler generates decent, optimized byte code. The following are some of the optimizations that the java compiler does.

Code Movement: moves simple expressions outside a loop if the variables used in that expression do not change inside the loop.

For Example:

int x = 15;
int y = 10;
int arr[10];
for (i=0; i<10;i++)

arr[i] = x + y;

Because x and y are invariant and do not change inside of the loop, their addition doesn't need to be performed for each loop iteration. So the compiler moves the addition of x and y outside the loop, thus creating a more efficient loop. For example, the optimized code could look like the following:

int x = 15;
int y = 10;
int z = x + y;
int arr[10];
for (i=0; i<10;i++)

arr[i] = z;

Constant folding: Constant folding refers to the compiler precalculating constant expressions. For example, examine the following code:

static final int length = 25;
static final int width = 10;
int res = length * width;

During execution these values will not be multiplied; instead, multiplication is done at compile time. The code for the following variable assignment is modified to produce bytecode that represents the product of width and length:

int res = 250;

Limited dead code elimination: No code is produced for statements like

if(false)

{
...
}

Dead code elimination does not affect the code's runtime execution. It does, however, reduce the size of the generated class file.

Other Compiler Related Options

There are also just in just-in-time (JIT) compilers that convert Java bytecodes into native code at run time. Several are available, and while they can increase the execution speed of your program dramatically, the level of optimization they can perform is constrained because optimization occurs at run time. A JIT compiler is more concerned with generating the code quickly than with generating the quickest code.

You can also use native code compilers that compile Java directly to native code should offer the greatest performance but at the cost of platform independence.

Viet Nam's Information Technology Industry

Trích từ bài viết đăng tải tại forum của trường Đại học Khoa Học Tự Nhiên TPHCM:

Thưa các anh chị,

Tôi có nhiều buổi nói chuyện với các bạn tôi và một số người về điều mà VN chúng ta đang đặt ra các con số. Con số 20,000 tiến sỹ rồi 1 triệu kỹ sư CNTT. Tất cả chỉ vì một mục đích 1 tỷ USD quá kém cõi.

Việc đặt ra 1 triệu kỹ sư CNTT và 20,000 tiến sỹ là con số khổng lồ, vượt quá khả năng hiện tại của VN trong ít nhất 15 năm tới. Con số này trở thành không tưởng nếu chúng ta nhìn lại về hệ thống đào tạo trên thế giới và VN. Hai khả năng quá xa vời này đang tạo nên một lỗ hỗng an ninh quá cao cho nguy cơ "Tiến sỹ giấy", phá hoại nội lực và kinh tế đất nước là rất có thể.

Và chúng ta lại có những báo cáo mang tính "nhầm lẫn" quá đáng khi nói rằng việc vượt qua 4 lần doanh thu về CNTT (500 triệu USD) so với năm 2003 (là 120 triệu) là con số đáng mừng? Chúng ta cần xem năm 2006 chúng ta làm được gì và năm 2007 chúng ta đã làm được gì và thực tế cho chúng ta thấy chúng ta đã suy yếu thế nào. Từ đây, hãy có một tầm nhìn khác cho một đột phá CNTT Việt Nam.

Phần mềm VN là "không có gì" ngoài cái vỏ bộc thô thiển trong "vấn nạn" nguỵ trang về "gia công phần mềm". Tôi ước tính con số để các anh chị thấy rằng, gia công hiện tại không phải có khả năng lợi nhuận.

1) Năm 2007, FCGV từ 650 lập trình viên giảm còn 450 lập trình viên. Điều tệ nhất là, hầu hết các lập trình viên có kinh nghiệm đề ra đi. Doanh thu công ty này không đạt chỉ tiêu và phải đối diện với việc "phá sản". Cuối cùng, FCG ở Mỹ quyết định bán cho CSC. Con số hiện nay của họ là 500 lập trình viên. Đi xuống!

2) PSD - do anh Thịnh làm chủ thì cho rằng năm 2007 là một năm tồi tệ nhất. Thua lỗ, cắt giảm nguồn nhân lực và "tìm cách để sống qua ngày" là chiến lược của CNTT.

3) TMA - Ông Nguyễn Hữu Lệ: Tạm bằng lòng với chỉ tiêu 750 kỹ sư. Trong khi đáng ra công ty này cần vượt qua con số 1200 kỹ sư. Lý do, tài chính và nguồn dự án bị đóng băng tại Bắc Mỹ (Canada)

4) Tân thiên niên Kỷ và IITS thì bị lỗ nặng và hai công ty với con số gần 100 lập trình viên buộc phải sát nhập lại nhau để cố giữ được tên công ty.

5) GlobalSoft: Một năm của những ảm đạm nhất chưa từng thấy.

6) FSoft: FS doanh thu 2006 là 18triệu USD, năm 2007 là hơn 29Triệu USD (10triệu USD lợi nhuận). Nhưng số nhân viên cũng tăng từ 1400 vào cuối năm 2006 lên 2500 vào cuối năm 2007. Tỷ lệ tăng doanh thu xấp xỉ bằng tỷ lệ tăng số nhân viên.Còn năng suất 1 người làm ra là: 10M USD ~ 160tỷ / khoảng 1800 nhân viên (Tính trung bình cả năm) ~ 90 triệu / người / năm => Với năng suất thấp như vậy thử hỏi phần mềm liệu có đáng được gọi là ngành công nghệ cao???

Nếu anh chị đi qua eTown 1,2 sẽ thấy sự trống rỗng còn đó của các building mà trước đó rất khó mà thuê.

Trả lời câu hỏi vì sao Harvey Nash không tuyển nhân viên vào lấp kín buidling mình (hiện nay tại eTown 2, công ty này đang có 75 lập trình viên trong khi họ có khả năng lấp đầy building với trên 150 lập trình viên) thì họ cho rằng, họ không muốn đổ vỡ theo bánh xe CNTT Việt Nam hiện tại.

Chúng ta đã quá tự hào vào cái đêm "Giải Sao Khuê". Tôi đáng buồn cho các nhà tổ chức đang vẽ ra một bức tranh mà họ chưa nhận thấy điều này. Có những công ty lỗ nặng nề cũng ôm trên tay chiếc cúp. Sao mà khó xem thế... sự dối trá sẽ là điều tàn ác nhất, phá huỷ kinh tế của một đất nước đang hồi sinh.

Chúng ta không nên đạt ra những con số để làm mục tiêu tăng trưởng. Con số 1 triệu kỷ sư CNTT để có doanh thu 1 tỷ USD là quá thấp về tính kinh tế trong khi khó tưởng với nguồn nhân lực VN.

Chúng ta đang khủng hoảng về gia công phần mềm vì chúng ta đang "làm mướn" mà thôi. Gia công VN đang chôn vùi tri thức trẻ sáng tạo, chôn vùi sự sáng tạo vốn đòi hỏi các bạn phải nỗ lực gấp nhiều lần. Gia công đang làm nhân lực VN xem rẽ giá trị chính mình để luyện giọng tiếng anh nói "bập bẹ" và ghép mình vào những "quy trình kiễu CMMI - 5 dối trá" làm tê liệt đi khả năng chủ động sáng tạo CN mà chính điều này mới là giá trị thực.

Tôi mong rằng, chính sách của Chính phủ VN sớm nhìn thấy một cách rõ ràng về giá trị nội lực và chiến lược sáng tạo mới là thúc đẫy hướng gia công. Chúng ta đang đi ngược lại tính tự nhiên và ép mình như việc ép con em chúng ta đọc thuộc lòng bài lịch sữ. Hãy cho tri thức trẽ những sáng tạo và đó là giá trị kinh tế thực sự mà tổ quốc cần.

Số liệu thống kê:
Việt Nam : Báo cáo toàn cảnh CNTT Việt Nam năm 2007
Ấn Độ: Indian Information Technology Industry

Laws of Software Evolution Revisited







Có sinh sản là có tiến hoá. Sinh vật có F1 F2, phần cứng có model1 model2, phần mềm có version1 version2. Sinh vật tiến hoá theo định luật Darwin, Mendel. Phần cứng tiến hoá theo định luật Moore, nano. Phần mềm tiến hoá theo định luật gì?

Giáo sư Lehman đã nghiên cứu vấn đề này từ tận những năm 1970 đến nay! Bài viết này giới thiệu 8 định luật ông phát hiện qua nghiên cứu rất nhiều phần mềm mã đóng. Phần mềm mã đóng có lịch sử từ khoảng 1950, mã mở, outsourcing mới chỉ ra đời mươi năm nay. Do đó để áp dụng cho 2 loại này, có lẽ phải mở rộng thêm 8 định luật này hoặc đưa ra những định luật khác chăng?

Định luật 1: Continuing Change

Phải chỉnh sửa phần mềm liên tục, nếu không mức độ hài lòng của khách hàng ngày càng giảm.

Có 2 lí do, liên quan đến lí thuyết về feedback control system của môn điều khiển học, vì rõ ràng khách hàng là người điều khiển lập trình viên:

  • Khả năng diễn đạt của khách hàng có hạn chế, họ cần x, nhưng lại nói thành y, lập trình viên nghe thành z
  • Nhu cầu của khác hàng thay đổi theo thời gian

Như vậy phần mềm phải tiến hoá liên tục nếu không muốn bị khách hàng xoá khỏi máy.

Định luật 2: Increasing Complexity

Khi tiến hoá, độ phức tạp của phần mềm luôn tăng, nếu không bỏ công sức để làm giảm nó xuống.

Đây chỉ là hệ quả của định luật 2 của môn nhiệt động học bao trùm mọi thứ trong vũ trụ, nói rằng entropy luôn tăng. Điều này có nghĩa chương trình ngày càng béo ra, cấu trúc ngày càng xấu tệ, cần phải refactor.

Định luật 3: Large Program Evolution

Các yếu tố liên quan đến quá trình tiến hoá (ý thích của khách hàng...) tuân theo qui luật phân phối xác suất chuẩn.

Định luật 4: Invariant Work-Rate

Độ ổn định là chỉ số quan trọng trong hệ thống điều khiển. Để đảm bảo độ tiến hoá ổn định, nhân sự cần ổn định qua thời gian.

Ai từng đọc quyển The mythical man-month đều biết: tăng thêm người vào team càng làm cho project đã chậm càng chậm hơn.

Định luật 5: Conservation of Familiarity

Ở mỗi phiên bản mới, phiên bản này chỉ thành công nếu những người liên quan (lập trình viên, nhân viên bán hàng, người dùng...) hiểu rõ sự khác biệt của phiên bản mới so với phiên bản trước. Do đó, phải bảo toàn hệ số góc của đường phát triển. Thay đổi nhanh quá sẽ bà con theo không kịp.

Định luật 6: Continuing Growth

Phải thêm tính năng vào phần mềm, nếu không mức độ hài lòng của khách hàng ngày càng giảm.

Có vẻ giống định luật 1, vì 2 cái nói về 2 hiện tượng khác nhau nhưng không phải không liên quan. Định luật 1 liên quan đến lí thuyết bất định Heisenberg của môn cơ học lượng tử: khách hàng không biết trước để có thể trình bày đầy đủ và chính xác mọi yêu cầu.

Ở định luật 6 thì ngược lại, khách hàng biết rõ họ cần 100 tính năng, nhưng do điều kiện tài chính, thời gian, tay nghề của lập trình viên... họ phải cắt bớt số tính năng xuống còn 60 để version 1 có thể hoàn thành kịp thời hạn. Sau đó, qua thời gian họ sẽ yêu cầu thêm tính năng còn thiếu vào version 2, 3.

Định luật 7: Declining Quality

Chất lượng của phần mềm càng ngày càng giảm nếu không được bảo trì và thay đổi cho phù hợp với điều kiện thực tế.

Theo thời gian mọi thứ đều tốt lên, nên về mặt tương quan, cái nào không tiến hoá sẽ tự động được coi là kém chất lượng. Ví dụ thời bao cấp mỗi tháng có nửa kí thịt thì được là có chất lượng sống cao, nhưng cũng nửa kí thịt đó (không thay đổi) thì hiện nay được coi là đói nghèo.

Định luật 8: Feedback System

Để có thể sửa chữa cải tiến, phải coi qui trình phát triển phần mềm là hệ thống điều khiển kiểu feedback.

Nguồn : Blog cộng đồng về CNTT
For more information: Laws of Software Evolution Revisited (1999)

Scheduling with Quartz

Batch solutions are ideal for processing that is time and/or state based:
  • Time-based: The business function executes on a recurring basis, running at pre-determined schedules.
  • State-based: The jobs will be run when the system reaches a specific state.
Batch processes are usually data-centric and are required to handle large volumes of data off-line without affecting your on-line systems. This nature of batch processing requires proper scheduling of jobs. Quartz is a full-featured, open source job scheduling system that can be integrated with, or used along side virtually any Java Enterprise of stand-alone application. The Quartz Scheduler includes many enterprise-class features, such as JTA transactions and clustering. The following is a list of features available:
  • Can run embedded within another free standing application
  • Can be instantiated within an application server (or servlet container).
  • Can participate in XA transactions, via the use of JobStoreCMT.
  • Can run as a stand-alone program (within its own Java Virtual Machine), to be used via RMI
  • Can be instantiated as a cluster of stand-alone programs (with load-balance and fail-over capabilities)
  • Supoprt for Fail-over
  • Support for Load balancing.
The following example demonstrates the use of Quartz scheduler from a stand-alone application. Follow these steps to setup the example, in Eclipse.
  1. Download the latest version of quartz from opensymphony.
  2. Make sure you have the following in your class path (project-properties->java build path):
    • The quartz jar file (quartz-1.6.0.jar).
    • Commons logging (commons-logging-1.0.4.jar)
    • Commons Collections (commons-collections-3.1.jar)
    • Add any server runtime to your classpath in eclipse. This is for including the Java transaction API used by Quartz. Alternatively, you can include the JTA class files in your classpath as follows
      1. Download the JTA classes zip file from the JTA download page.
      2. Extract the files in the zip file to a subdirectory of your project in Eclipse.
      3. Add the directory to your Java Build Path in the project->preferences, as a class directory.
  3. Implement a Quartz Job: A quartz job is the task that will run at the scheduled time.
    public class SimpleJob implements Job {
    public void execute(JobExecutionContext ctx) throws JobExecutionException {
    System.out.println("Executing at: " + Calendar.getInstance().getTime() + " triggered by: " + ctx.getTrigger().getName());
    }
    }

  4. The following piece of code can be used to run the job using a scheduler.
    public class QuartzTest {
    public static void main(String[] args) {
    try {
    // Get a scheduler instance.
    SchedulerFactory schedulerFactory = new StdSchedulerFactory();
    Scheduler scheduler = schedulerFactory.getScheduler();

    long ctime = System.currentTimeMillis();

    // Create a trigger.
    JobDetail jobDetail = new JobDetail("Job Detail", "jGroup", SimpleJob.class);
    SimpleTrigger simpleTrigger = new SimpleTrigger("My Trigger", "tGroup");
    simpleTrigger.setStartTime(new Date(ctime));

    // Set the time interval and number of repeats.
    simpleTrigger.setRepeatInterval(100);
    simpleTrigger.setRepeatCount(10);

    // Add trigger and job to Scheduler.
    scheduler.scheduleJob(jobDetail, simpleTrigger);

    // Start the job.
    scheduler.start();
    } catch (SchedulerException ex) {
    ex.printStackTrace();
    }
    }
    }

    A trigger is used to define the schedule in which to run the job.
For more information on batch processing visit: "High volume transaction processing in J2EE"

Java EE 6 Highlights

The key features of Java EE 6 (Java Enterprise Edition version 6) are:

Modular Platform - Java EE 6 introduces profiles targeted for particular segment of users like web developers or mobile developers. Java Profiles allows you to select Java EE 6 features to be included in a profile. This allows creating smaller runtime with only the modules and extensions you need.

Extensibility - Scripting languages and extensions are now treated as “first class citizens” and can be easily integrated with the core platform. Third party libraries will be able to self-register.

Annotations across Web API - No more manual editing of web.xml (yeah!).

RESTful web services - Java EE 6 will support creating RESTful web services out of the box.

For more information: Introduction to Java 6.0 New Features