www.ijraset.com
Vol. 1 Issue V, December 2013 ISSN: 2321-9653
I N T E R N A T I O N A L J O U R N A L F O R R E S E A R C H I N A P P L I E D S C I E N C E AN D E N G I N E E R I N G T E C H N O L O G Y (I J R A S E T)
Reducing Error Signal in Multilayer Perceptron Neural Networks using MLP for Label Ranking Kalyana Chakravarthy Dunuku1
V.Saritha2
HOD1, Department Of CSE, Sri Venkateswara Engineering College, Piplikhera, Sonepat, Haryana, Pin-131039
Assoc.Prof.2, Department of CSE Sri Kavita Engineering College, Karepalli, Khammam, A.P. Pin- 507122
Abstract: - This paper describes a simple tactile probe for identifying error signal
in Multilayer. In multilayer having the
number of hidden layers error signal can be process as irrespective manner so difficult to find out the error signal. The multilayer perceptron having the number of hidden layers with one output layer. This networks are fully connected i.e. a neuron in any layer of this network is connected to all the nodes/neurons in the previous layer signal flow through the network progress in a forward direction from left to right and on a layer by layer. In this networks we can identify the two kinds of networks. First one is Function Signal-A function signal is an input signal that comes in at the Input end of the network. Second one is Error Signal- an error signal originates at an output neuron of the network and propagates backward i.e. layer by layer through the network. In this paper, we adapt a multilayer perceptron algorithm for label ranking. We focus on the adaptation of the BackPropagation (BP) mechanism.
Keywords: Label Ranking, back-propagation, multilayer perceptron.
1.
INTRODUCTION
perceptron that has multiple layers. Rather, it contains many
This class of networks consists of multiple layers of
perceptrons that are organized into layers, leading some to
computational units, usually interconnected in a feed-forward
believe that a more fitting term might therefore be "multilayer
way. Each neuron in one layer has directed connections to the
perceptron network". Moreover, these "perceptrons" are not
neurons of the subsequent layer [11][18]. In many applications
really perceptrons in the strictest possible sense, as true
the units of these networks apply a sigmoid function as an
perceptrons are a special case of artificial neurons that use a
activation function.
threshold activation function such as the Heaviside step
Multilayer Perceptron. The term "multilayer perceptron" often
function, whereas the artificial neurons in a multilayer
causes confusion. It is argued the model is not a single
perceptron are free to take on any arbitrary activation
Page 40
www.ijraset.com
Vol. 1 Issue V, December 2013 ISSN: 2321-9653
I N T E R N A T I O N A L J O U R N A L F O R R E S E A R C H I N A P P L I E D S C I E N C E AN D E N G I N E E R I N G T E C H N O L O G Y (I J R A S E T) function[16][18]. Consequently, whereas a true perceptron performs binary classification, a neuron in a multilayer perceptron is free to either perform classification or regression, depending upon its activation function. The two arguments raised above can be reconciled with the name "multilayer perceptron" if "perceptron" is simply interpreted to mean a binary classifier, independent of the specific mechanistic implementation of a classical perceptron. In this case, the entire network can indeed be considered to be a binary classifier with multiple layers[11][12]. Furthermore, the term
Fig 1. Artificial neural network, Three layers MLP
"multilayer perceptron" now does not specify the nature of the layers; the layers are free to be composed of general artificial neurons, and not perceptrons specifically. This interpretation of the term "multilayer perceptron" avoids the loosening of the definition of "perceptron" to mean an artificial neuron in general. 1.1. ARCHITECTURE The network topology used in this study is based on fully connected feed-forward ANNs. The number of nodes in the input layer is equal to the number of features presented by the data, while the number of nodes in the output layer (L) is equal to the number of classes that this data map to. At least one hidden layer must be added to the architecture in order to treat the non-linear separation among classes. Several networks with one and two hidden layers, with different number of nodes in each hidden layer, have been used[11][12][15]. The architecture for a multilayer perceptron with two hidden layers.
The fig.2. Depicts a portion of the multilayer perceptron. Two kinds of signals are identifies in this network.
Page 41
www.ijraset.com
Vol. 1 Issue V, December 2013 ISSN: 2321-9653
I N T E R N A T I O N A L J O U R N A L F O R R E S E A R C H I N A P P L I E D S C I E N C E AN D E N G I N E E R I N G T E C H N O L O G Y (I J R A S E T)
i). Function signal
2.
A function signal is an input signal tbat comes in at the
The computation of an estimate of the gradient vector.
Gradient of
error surface with respect to the weights
input end of the network, propagates forward through the
connected to the inputs of a neuron. Which is needed for the
network, and emerges at the out put end if the network as ab
backward pass through the network.
output signal[11]. We refer to such a signal as a function signal In many real-world applications, assigning a single label
for two reasons. First, it is presumed to perform a useful function at the output of the network. Second, at each neuron of the network through which a function signal passes, the signal is calculated as a function of the input and associated weights applied to that neuron. The function signal is also referred to as
to an example is not enough. For instance, when trading in the stock market based on recommendations from financial analysts, predicting who is the best analyst does not suffice because 1) he/she may not make a recommendation in the near future and 2)
the input signal.
we
may prefer to take into account
recommendations of multiple analysts, to be on the safe ii). Error signal. An error signal originates at an output neuron of the
side[1][2][3][5]. Hence, to support this approach, a model
network, and propagates backward(layer by layer) through the
single one. Such a situation can be modeled as a Label Ranking
network. We refer to it as an error signal because its
(LR) problem: a form of preference learning, aiming to predict a
computation by every neuron of the network involves an error-
mapping from examples to rankings of a finite set of labels.
dependent function in one form or another[11]. The output
Recently, quite some solutions have been proposed for the label
neurons constitute the output layers of the network. The
ranking problem. including one based on the Multilayer
remaining neurons constitute hidden layer of the network. Thus
Perceptron algorithm (MLP). MLP is a type of neural network
the hidden units are not part of the output or input of the
architecture, which has been applied in a supervised learning
network hence their designation as “hidden�[15] The first
context using the error back-propagation (BP) learning
hidden layer is fed from the input layer made up of sensory
algorithm. In this paper, we try a different approach to the
units, the resulting outputs of the first hidden layer are in turn
simple adaptation proposed earlier[1][8][9]. We adapt the BP
applied to the next hidden layer and so on for the rest of the
learning mechanism to LR. More specifically, we investigate
network[11][12][15]. Each hidden or output neuron of
how the error signal explored by BP can use information from
a
multilayer perceprton is designed to perform two computations: 1.
The computation of the function signal appearing at the
output of a neuron, which is expressed as a continuous nonlinear function of the input signal and synaptic weights
should predict a ranking of analysts rather than suggesting a
the LR loss function. We introduce six approaches and evaluate their (relative) performance. We also show some preliminary experimental results that indicate whether our new method could compete with state-of-the-art LR methods.
associated with that neuron.
Page 42
www.ijraset.com
Vol. 1 Issue V, December 2013 ISSN: 2321-9653
I N T E R N A T I O N A L J O U R N A L F O R R E S E A R C H I N A P P L I E D S C I E N C E AN D E N G I N E E R I N G T E C H N O L O G Y (I J R A S E T) 2. PRELIMINARIES
learning algorithm rather than the network itself. The network
Throughout this paper, we assume a training set T ={xn, πn} consisting oft examples xn and their associated label
used isgenerally of the simple type shown in figure 4 [11][16][17][18].
rankings πn. Such a ranking is a permutation of a finite set of
A Back Propagation network learns by example. You give
labels L={λ1,...,λk}, given k, taken from the permutation space
the algorithm examples of what you want the network to do and
ΩL.Each example xn consists of m attributes xn = {a1,...,am}and
it changes the network’s weights so that, when training is
is taken from the example space X. The position of λa in a
finished, it will give you the required output for a particular
rankingπnis denoted by πn(a) and assumes a value in the
input. Back Propagation networks are ideal for simple Pattern
set{1,...,k}
Recognition and Mapping Tasks 4[12][15]. As just mentioned,
2.1
Back-Propagation Algorithm.
to train the network you need to give it examples of what you
The Back Propagation network to be the quintessential Neural Net. Actually, Back Propagation is the training or
want – the output you want (called the Target) for a particular input as shown in Figure 5.
represented by 1 and a white by 0 as in the previous examples). Fig. 5. Back Propagation training set.
The input and its corresponding target are called a Training Pair.
So, if we put in the first pattern to the network, we would like the output to be 0 1 as shown in figure 6. (a black pixel is
Page 43
www.ijraset.com
Vol. 1 Issue V, December 2013 ISSN: 2321-9653
I N T E R N A T I O N A L J O U R N A L F O R R E S E A R C H I N A P P L I E D S C I E N C E AN D E N G I N E E R I N G T E C H N O L O G Y (I J R A S E T)
Fig. 6. applying a training pair to a network The network is first initialised by setting up all its weights to be small random numbers – say between –1 and +1. Next, the input pattern is applied and the output calculated (this is called the forward pass). The calculation gives an output which is completely different to what you want (the Target), since all the weights are random. We then calculate the Error of each neuron, which is essentially: Target - Actual Output (i.e. What you want – What you actually get). This error is then used mathematically to change the weights in such a way that the
Fig. 7. single connection learning in a Back Propagation
error will get smaller. In other words, the Output of each neuron
network.
will get closer to its Target(this part is called the reverse pass).
The connection we’re interested in is between neuron A
The process is repeated again and again until the error is
(a hidden layer neuron) and neuron B (an output neuron)and has
minimal Let's do an example with an actual network to see how
the weight WAB. The diagram also shows another connection,
the process works[11][15]. We’ll just look at one connection
between neuron A and C, but we’ll return to that later[10][14].
initially, between a neuron in the output layer and one in the
2.2. The algorithm works :
hidden layer in fig. 7. Step 1. First apply the inputs to the network and work out the output – remember this initial output could be anything, as the initial weights were random numbers.
Page 44
www.ijraset.com
Vol. 1 Issue V, December 2013 ISSN: 2321-9653
I N T E R N A T I O N A L J O U R N A L F O R R E S E A R C H I N A P P L I E D S C I E N C E AN D E N G I N E E R I N G T E C H N O L O G Y (I J R A S E T)
Step 2. Next work out the error for neuron B. The error is What
let’s clear that up by explicitly showing all the calculations
for a full sized network with 2 inputs, 3 hidden layer neurons
you want – What you actually get, in other words:
and 2 output neurons as shown in figure. 8. W+ represents the ErrorB = OutputB(1-OutputB)(TargetB– OutputB)
new, recalculated, weight, whereas W represents the old
The “Output(1-Output)” term is necessary in the
weight[16][17][18].
equation because of the Sigmoid Function – if we were only using a threshold neuron it would just be (Target –Output). Step 3. Change the weight. Let W+AB be the new (trained) weight and WAB be the initial weight. W+AB= WAB + (ErrorBx OutputA).Notice that it is the output of the connecting neuron (neuron A) we use (not B). We update all the weights in the output layer in this way. Step 4. Unlike
Calculate the Errors for the hidden layer neurons. the
output
layer
we
can’t
calculate
these Fig. 8 Three layers full sized network
directly(because we don’t have a Target), so we Back Propagate them from the output layer (hence the name of the algorithm).
All the calculations for a reverse pass of Back Propagation.
This is done by taking the Errors from the output neurons and
1. Calculate errors of output neurons
running them back through the weights to get the hidden layer
δα = outα(1 - outα) (Targetα- outα)
errors. For example if neuron A is connected as shown to B and
δβ = outβ(1 - outβ) (Targetβ- outβ)
C then we take the errors from B and C to generate an error for
2. Change output layer weights
A.
W+Aα = WAα+ ηδα outA
W+Aβ = WAβ+ ηδβ outA
W+Bα = WBα+ ηδα outB
W+Bβ = WBβ+ ηδβ outB
W+Cα = WCα+ ηδα outC
W+Cβ = WCβ+ ηδβ outC
ErrorA= OutputA(1 - OutputA)(ErrorBWAB + ErrorCWAC) Again, the factor “Output (1 - Output )” is present
3. Calculate (back-propagate) hidden layer errors
because of the sigmoid squashing function.
δA = outA(1 – outA) (δαWAα + δβWAβ) Step 5. Having obtained the Error for the hidden layer neurons
δB = outB(1 – outB) (δαWBα + δβWBβ)
now proceed as in Step 3 to change the hidden layer weights. By
δC = outC(1 – outC) (δαWCα + δβWCβ)
repeating this method we can train a network of any number of
4. Change hidden layer weights
layers.
W+λA = WλA + ηδA inλ
W+ΩA = W+ΩA+ ηδA inΩ
2.3. Calculation of Reverse pass of Back Propagation.
W+λB = WλB + ηδB inλ
W+ΩB = W+ΩB+ ηδB inΩ
Page 45
www.ijraset.com
Vol. 1 Issue V, December 2013 ISSN: 2321-9653
I N T E R N A T I O N A L J O U R N A L F O R R E S E A R C H I N A P P L I E D S C I E N C E AN D E N G I N E E R I N G T E C H N O L O G Y (I J R A S E T)
W+λC = WλC + ηδC inλ
W+ΩC = W+ΩC+ ηδC inΩ
δ1 = δx w1 = -0.0406 x 0.272392 x (1-o)o = -2.406x10-3
The constant η (called the learning rate, and nominally equal to one) is put in to speed up or slow down the learning if
δ2= δx w2 = -0.0406 x 0.87305 x (1-o)o = -7.916x10-3
required. 2.4. Example:
New hidden layer weights: w3+=0.1 + (-2.406 x 10-3x 0.35) = 0.09916.
Consider the simple network below:
w4+= 0.8 + (-2.406 x 10-3x 0.9) = 0.7978. w5+= 0.4 + (-7.916 x 10-3x 0.35) = 0.3972. w6+= 0.6 + (-7.916 x 10-3x 0.9) = 0.5928. (iii) Old error was -0.19. New error is -0.18205. Therefore error has reduced. 2.5.
Total Error in Network: Error falls to some pre-determined low target value and
Assume that the neurons have a Sigmoid activation function and
then it stops.
(i) Perform a forward pass on the network. (ii) Perform a reverse pass (training) once (target = 0.5). (iii) Perform a further forward pass and comment on the result. Answer: (i) Input to top neuron = (0.35x0.1)+(0.9x0.8) = 0.755. Out = 0.68. Input to bottom neuron = (0.9x0.6)+(0.35x0.4) = 0.68. Out = 0.6637. Input to final neuron =(0.3x0.68)+(0.9x0.6637) = 0.80133.Out =0.69. (ii) Output error δ=(t-o)(1-o)o = (0.5-0.69)(1-0.69)0.69 = -0.0406. New weights for output layer w1+= w1+(δx input) = 0.3 + (-0.0406x0.68) = 0.272392. +
w2 = w2+(δx input) = 0.9 + (-0.0406x0.6637) = 0.87305.
Fig. 9. The network keeps training all the patterns repeatedly until the total
Errors for hidden layers:
Page 46
www.ijraset.com
Vol. 1 Issue V, December 2013 ISSN: 2321-9653
I N T E R N A T I O N A L J O U R N A L F O R R E S E A R C H I N A P P L I E D S C I E N C E AN D E N G I N E E R I N G T E C H N O L O G Y (I J R A S E T) Note that when calculating the final error used to stop the
This stops the network overtraining. It does this by having a
network (which is the sum of all the individual neuron errors for
second set of patterns which are noisy versions of the training
each pattern) you need to make all errors positive so that they
set. Each time after the network has trained; this set (called the
add up and do not subtract[11][16][17][18].
Validation Set) is used to calculate an error. When the error
Once the network has been trained, it should be able to
becomes low the network stops.
recognise not just the perfect patterns, but also corrupted In fact
The following shows the use of validation set.
if we deliberately add some noisy versions of the patterns into
When the network has fully trained, the Validation Set error
the training set as we train the network (say one in five), we can
reaches a minimum. When the network is overtraining
improve the network’s performance in this respect. The training
(becoming too accurate) the validation set error starts rising [7].
may also benefit from applying the patterns in a random order to
If the network overtrains, it won’t be able to handle noisy data
the network. There is a better way of working out when to stop
so well.
network training - which is to use a Validation Set[11][13][16].
Fig.10 Use of validation sets 3. MULTILAYER PERCEPTRON FOR LABEL RANKING
The tricky point of adapting an MLP for LR is the weight
Our adaptation of MLP for LR essentially consists of 1) the
corrections in the BP process: minimizing the individual errors
method to generate a ranking from the output layer and 2) the
does not necessarily lead to minimizing the LR loss. We
error functions guiding the BP learning process. The output
propose six approaches to define the error signal cj at the output
layer contains k neurons (one for each label). The output yj of a
layer[5][6]. The weight connection wji(n) is updated based on
neuron j at the output layer does not represent a target value or
the estimated cj(n) using the delta rule Δwji(n)=ηcj(n)yi(n).
class but rather the score associated with a label λj. By ordering all the scores, the predicted ranks π’n(j) of the label λj and, thus,
Local Approach (LA). The error signal is the individual error of each output neuron,
the predicted ranking[1][2].
Page 47
www.ijraset.com
Vol. 1 Issue V, December 2013 ISSN: 2321-9653
I N T E R N A T I O N A L J O U R N A L F O R R E S E A R C H I N A P P L I E D S C I E N C E AN D E N G I N E E R I N G T E C H N O L O G Y (I J R A S E T) cj(n)=ej(n)=πn(j)−π’n(j), as in the original MLP. The LR
error, eτ, is only used to evaluate the activation of the BP.
output layer are important to define the error signal and are considered independently of the neurons they connect to. The
Global Approach (GA).The error signal is defined in terms of the LR error. In this case, it is simply given by cj(n)=eτ(n)
error signal denoted cji(n) is associated with the weight of the connection between output neuron i and hidden neuron j. This is
Combined Approach (CA). CA is a combination between
similar to WSGA but we rank all weight connections
GA and LA, cj(n)= ej(n)eτ(n). We note that a neuron which
individually, rather than the average weights for each output
returns the correct position πn(j)=π’n(j) (i.e.,ej(n) = 0) is not
neuron. The weight corrections are given by Δwji(n)=ηcji(n)yi(n),
penalized even if eτ >0.
where:
Weight-Based Signed Global Approach (WSGA).The error signal is defined in terms of the LR error and the incoming
−eτ(n),
if
pgw(ji) >
eτ(n),
if
pgw(ji) ≤
cji(n) =
weight connections of the output layer. We assume that a high LR error means that some weights of neurons are too high and
4.
other are too low. The output neurons are ranked according to their average weights ῶj = ∑
= ⍵ij resulting in a position
p⍵(j)∈[1,...,k]. The error of the neurons with a position above the mean is negative and it is positive otherwise: −eτ(n)
EXPERIMENTAL RESULTS
The goal is to compare the performance of the proposed approaches on different datasets. The datasets used for the evaluation. These datasets, which are commonly used for LR, are presented in Table 1. Our approach starts
if pw(j) > ( + 0.5 )
Our approach starts by normalizing all attributes, and separating the dataset into a training and a test set. On each
Cj =
eτ(n)
if pw(j) < ( + 0.5 )
0
if pw(j) = ( + 0.5)
dataset we tested the six approaches with h= 3 hidden neurons, η=0.2, using 5 epochs with 5 random restarts. The error
(1)
estimation methodology is 10-fold cross-validation. The results Score-Based
Signed
Global
Approach
(SSGA).The
motivation for SSGA is the same as for WSGA. The difference
are presented in terms of the similarity between the rankings πi and π’i with the Kendall τ coefficient.
is that we rank the output neuron scores yj instead of the input
In Table 2, we show the resulting τ-values for each
weights. The positions of the weights, pw(j) is replaced in eq. 1
approach, and associated rank (lower is better) per dataset. The
with the positions of the scores, ps(j)
bottom row shows the average rank for each approach, which
Individual
Weight-Based
Signed
Global
Approach
(IWSGA). This assumes that all the weight connections at the
allows us to compare the relative performance of the approaches using the Friedman test with post-hoc Nemenyi test [13].
Table 1. Datasets for LR
Page 48
www.ijraset.com
Vol. 1 Issue V, December 2013 ISSN: 2321-9653
I N T E R N A T I O N A L J O U R N A L F O R R E S E A R C H I N A P P L I E D S C I E N C E AN D E N G I N E E R I N G T E C H N O L O G Y (I J R A S E T)
Table 2.Experimental results of MLP-LR and their ranks
implies that for each pair of approaches Ai and Aj, if Ri < The Friedman test proves that the average ranks are
Rj−CD, then Ai is significantly better than Aj. Hence we can see
significantly unequal (with α=1% ).Then the Nemenyi test gives
from the table that approaches LA and CA significantly
us a critical difference of CD=2.225 (with α=1%). The test
outperform all other approaches except for SSGA. However, at
Page 49
www.ijraset.com
Vol. 1 Issue V, December 2013 ISSN: 2321-9653
I N T E R N A T I O N A L J O U R N A L F O R R E S E A R C H I N A P P L I E D S C I E N C E AN D E N G I N E E R I N G T E C H N O L O G Y (I J R A S E T)
α= 10% the critical difference becomes CD=1.712, so at this
can substantially improve the results when tweaking the number
significance level CA significantly outperforms SSGA too.
of epochs. For some approaches using more epochs is better, but
These experiments are performed with a rather arbitrary set
for others this monotonicity does not hold. We see similar
of parameters. Varying parameters such as the number of hidden
behavior when varying the number of stages and hidden
neurons in the MLP, the number of epochs used when learning
neurons. When the dataset at hand has relatively many
the neural network, and the number of random restarts, could
attributes, our approaches have relatively many input signals in
benefit performance. To illustrate this, Figure 1b displays the
the MLP.
variation of τ-values for the different approaches on the Iris dataset, when varying the number of epochs. As we can see, we
Fig. 9 Results of Kendall’s τ correlation coefficient.
Hence there are many more connections with the hidden
incorporate the individual errors perform significantly better
layer, and much more interactions between the neurons in the
than the methods that focus on the LR error. However, the best
network.
results are obtained by combining both errors (CA). A 5.
CONCLUSIONS
comparison
with
results
published
for
other
methods
In this paper, the most commonly used networks consist of
additionally indicates that our method has the potential to
an input layer, a single hidden layer and an output layer. The
compete with other methods. This holds even though no
input layer size is set by the type of pattern or input you want
parameter tuning was carried out, which is known to be essential
the network to process. And reducing the error signal in
for learning accurate networks. Our method becomes more
multilayer perceptron neural network shown in example
competitive when the data contains more attributes; this
Empirical results indicate that the two methods that directly
increases the amount of input neurons, and the MLP-LR
Page 50
www.ijraset.com
Vol. 1 Issue V, December 2013 ISSN: 2321-9653
I N T E R N A T I O N A L J O U R N A L F O R R E S E A R C H I N A P P L I E D S C I E N C E AN D E N G I N E E R I N G T E C H N O L O G Y (I J R A S E T)
predictions benefit from the more complex network. As future
[9]. Cheng, W., Huhn, J.C., H¨ ullermeier, E.: Decision
work, apart from parameter tuning we will investigate other
tree and instance-based learning for label ranking. In: ICML
ways of combining the local and global errors and we will
(2009)
investigate how to give more importance to higher ranks. 6.
[10]. de S´ a, C.R., Soares, C., Jorge, A.M., Azevedo, P.,
REFERENCES
Costa, J.: Mining Association Rules for Label Ranking. In:
[1]. Geraldina Ribeiro, Wouter Duivesteijn, Carlos
Huang, J.Z., Cao, L., Srivastava, J. (eds.) PAKDD 2011, Part
Soares, and Arno Knobbe Multilayer Perceptron for Label
II. LNCS, vol. 6635, pp. 432–443. Springer, Heidelberg
Ranking. (2011)
(2011)
[2]. Aiguzhinov, A., Soares, C., Serra, A.P.: A SimilarityBased Adaptation of Naïve Bayes for Label Ranking:
[11]. Haykin, S.: Neural Networks: a comprehensive foundation, 2nd edn (1998).
Application to the Metalearning Problem of Algorithm Recommendation. In: Discovery Science (2010).
[12]. F. J. Maldonado Williams-Pyro, M. T. Manry University of Texas at Arlington Electrical Engineering Dept
[3]. Vembu, S., G¨ artner, T.: Label Ranking Algorithms: A Survey. In: F¨ urnkranz, J., H¨ullermeier, E. (eds.)
Optimal Pruning of Feedforward Neural Networks Based upon the Schmidt Procedure.
Preference Learning. Springer (2010)
[13] K. Liu, S. Subbarayan, R.R.Shoults, M.T. Manry,
[4]. H¨ ullermeier, E., F¨urnkranz, J.: On loss functions in
C.Kwan,
F.L. Lewis, J.Naccarino, "Comparison of Very
label ranking and risk minimization by pairwise learning.
Short-Term Load Forecasting Techniques,"IEEE Transactions
JCSS 76(1), 49–62 (2010)
on Power Systems, vol.11, no.2, May 1996, pp. 877-882.
[5]. Brinker, K., H¨ ullermeier, E.: Label Ranking in
[14] V. Maniezzo, ìGenetic evolution of the topology and
Case-Based Reasoning. In: Weber, R.O., Richter, M.M. (eds.)
weight distribution of Neural Networksî, IEEE Transaction on
ICCBR 2007. LNCS (LNAI), vol. 4626, pp. 77–91. Springer,
Neural Networks, 1994, vol. 5, No. 1, pp. 39-53.
Heidelberg (2007).
[15] Ponnapalli, ìA formal selection and pruning
[6]. Dekel, O., Manning, C.D., Singer, Y.: Log-linear
algorithm for feedforward artificial network optimizationî,
models for label ranking. In: Advances in Neural Information
IEEE Transaction on Neural Networks, 1999, vol. 10, No. 4,
Processing Systems (2003)
pp. 964-968.
[7]. H¨ ulermeier, E., F¨urnkranz, J., Cheng, W., Brinker,
[16]. Alsmadi, M. K. S., Omar, K., & Noah, S. A. (2009).
K.: Label ranking by learning pairwise preferences. Artif.
Back Propagation Algorithm: The Best Algorithm Among the
Intell., 1897–1916 (2008)
Multi-layer Perceptron Algorithm. International Journal of
[8]. Cheng, W., Dembczynski, K., H¨ ullermeier, E.:
Computer Science and Network Security, 9 (4), pp. 378-383.
Label Ranking Methods based on the Plackett-Luce Model. In: ICML (2010)
[17]. Bi, W., Wang, X., Tang, Z., & Tamura, H. (2005). Avoiding the Local Minima Problem in Backpropagation
Page 51
www.ijraset.com
Vol. 1 Issue V, December 2013 ISSN: 2321-9653
I N T E R N A T I O N A L J O U R N A L F O R R E S E A R C H I N A P P L I E D S C I E N C E AN D E N G I N E E R I N G T E C H N O L O G Y (I J R A S E T) Algorithm with Modified Error Function. IEICE
Trans.
Fundam. Electron. Commun. Comput. Sci., E88-A (12), pp. 3645-3653. [18]. Nazri Mohd Nawia, R.S. Ransingb, Mohd Najib Mohd Sallehc, Rozaida Ghazalid, Norhamreeza Abdul Hamid The Effect Of Gain Variation In Improving Learning Speed Of Back
Propagation
Neural
Network
Algorithm
On
Classification Problems p 120-124.
Page 52